The Cognitive Corner

The Myth of Personalization: Why education keeps outsourcing the work of Teaching

Rebecca A. Huggins

If Winston Smith were real, and not a figment of George Orwell’s 1984, he would undoubtedly be wary of any education or civic institution that outsourced epistemic authority to platforms. Seventy-seven years after the publication of Orwell’s dystopia, the worry looks less like fiction: high-profile launches, like the First Lady’s public appearance alongside her own humanoid robot for her Fostering the Future Together initiative, tout technological spectacle as educational reform. 

John Steinbeck once claimed humanity’s greatest challenge would be: “success, plenty, comfort, and ever-increasing leisure.” Have we arrived at an “anesthetic of satisfaction” that breeds apathy, erodes collective resilience, and steers us into self-destructive traps? 

This may seem fatalistic to some, but when we outsource knowledge, learning, and educating our children to tech-based platforms, what are we gaining in return? If one thing has united the push for tailored learning movements over the past forty years, it has been a systemic disconnect between the optimistic, often revolutionary promises of educational innovation and the modest, inconsistent results found when those reforms are implemented at scale. The clichéd mantra of “meeting students where they are through personalized learning” is what drives many of these educative reforms. It has that illustrious, seemingly elusive glow of equity. But equity is achievable when districts invest in strong instruction and teacher development, not when they outsource teaching to technology.

Benjamin Bloom’s “two-sigma claim”—that one-to-one tutoring could raise student achievement by two standard deviations—has become ubiquitous in debates about personalization. While Bloom’s theory is famous, it was built on a very specific and narrow foundation that has proven difficult to replicate at scale. His iconic graph showing the massive benefit of tutoring was, in fact, only illustrative; hand-drawn, in a smooth, stylized fashion, and not actually fitted to any real data (von Hippel, 2024). The actual studies he relied on were three-week experiments on narrow unfamiliar topics and were based on dissertation studies by his PhD students, Joanne Anania and Arthur J. Burke (von Hippel, 2024). The studies themselves relied on specialized tests on topics like cartography and probability, where students started with zero knowledge and were assessed on tests the researchers themselves created. Yet when tutoring is measured against broader, standardized outcomes, the effect size shrinks considerably: estimates often fall to roughly 0.27 to 0.37 standard deviations (von Hippel, 2024). That gap between the legend and the evidence matters because it underwrites a policy impulse to scale tutoring-style solutions without the contextual supports that make small experiments work. 

Modern researchers largely agree that there is a significant gap in Bloom’s famous hand-drawn graph, and the actual data it purportedly represented. But this doesn’t mean that it has stopped advocates from framing “tailored learning” as a moral necessity to escape a system based on averages. Ironically, the moral necessity here is recognizing the glaring evidence gap in much of the research that bolsters these initiatives. Advocates have presented tailored learning as the “golden ticket” to all the equity concerns that traditional classrooms allegedly are unable to address. Whether it’s through flipped classroom initiatives, mastery-based learning models, personalized learning pathways, or competency-based education (CBE), tailored learning focuses on the idea that if we can control the pace and focus of instruction to the specific needs and goals of each student, we can effectively erase achievement gaps. While the empirical foundation for tailored learning is remarkably thin, often relying on trivial to small effect sizes, we have, nevertheless, continued to offload the art of teaching to technology. Take, for instance, flipped classrooms that cropped up in the early 2000s. Much like the ideologies that have perpetuated tailored learning, flipped classrooms were originally promoted as an effective way to blend learning—where the “tailoring” occurs by shifting the primary delivery of content (often via online video) to a student’s own time and pace. Despite their popularity, flipped classrooms lack a robust, conclusive research base. A meta-analysis of 115 effect-size comparisons found minimal effect (g=0.193), and many studies consisted predominantly of conference papers or proceedings, rather than rigorous peer-reviewed journal articles (Cheng et al., 2019). 

Similarly, while personalized learning (PL) seems promising in theory, there are very few evaluations of students’ learning outcomes in such programs. A major study of highly personalized schools found an average gain of only 3 percentile points in math, while reading estimates were near zero and statistically insignificant (Pane et al., 2017). A 3‑point percentile bump is far too small to erase a 3–4‑year achievement lag: because percentiles are relative ranks rather than linear measures of learning, that size of gain would likely take many years—often a decade or more—to eliminate unless it were accompanied by much larger, sustained scale‑score growth and intensive remediation. By the time such incremental gains accumulate, the student would likely be finishing secondary school and be preparing for college or the workforce, long past the point where a brief tech intervention could have changed their trajectory. Even more alarming, the modest average gains in these studies mask wide variation: some schools showed large negative effects on student achievement (Pane et al., 2017).  In other words, personalized learning can, in many instances, do more harm than good. 

Likewise, Competency-Based Education (CBE), the latest iteration of these tailored learning models, is often portrayed as a bridge between higher education and the labor market, but a systematic literature review of K-12 CBE implementation from 2000 to 2019 revealed that the promise of CBE far outpaces the data (Evans, 2020). Almost two-thirds of the research reviewed was qualitative, and based on small, unrepresentative samples, which limits the ability to generalize findings across different contexts. Perhaps most importantly, CBE does little to close achievement gaps; if anything, it disservices low-achieving students who are far less likely than their high-achieving peers to take ownership of the “reassessment and recovery” opportunities that are central to the CBE model (Evans, 2020). Many competency models have and do atomize learning into a supermarket shelf of unrelated tasks, removing the subjective articulation and socialization that a human educator provides (Ramírez Naranjo, 2022). Student success hinges on the quality of instruction, which technical personalization often fragments.  

Finally, AI tutoring represents the newest frontier in the tailored-learning movement. Frequently marketed as the ultimate solution to the “two-sigma problem,” high-profile innovators like Sal Khan launched AI tools like Khanmigo, arguing that AI can make one-to-one tutoring affordable at scale (von Hippel, 2024). Yet despite Sal Khan’s promise that the AI tutor would be the “biggest positive transformation that education has ever seen,” internal leadership now admits that for many students, the tool was a “non-event” and they simply “didn’t use it much” (Meyer, 2026). Despite WestEd’s randomized trial which reported some small improvements in algebra-readiness after one semester in schools using Khanmigo (Barnett, 2026), Meyer (2026) points to the inflated effect sizes, noting that its strongest claimed effects were in fact achieved by excluding 95% of the study population. It would seem that, overall, most chatbots lack the sensitivity and relationship-building capabilities required for effective tutoring (Meyer, 2026). 

Another model, Stanford’s Tutor CoPilot, uses AI to surface real-time, expert-like guidance for novice tutors, suggesting strategies drawn from the latent reasoning of experienced educators (Wang et al., 2024). That study points to a model in which technology amplifies human expertise rather than replaces it. Still, if we prioritized strengthening teachers’ content knowledge and pedagogical reasoning in preparation and professional learning, novice tutors and classroom teachers alike would be less dependent on AI scaffolds to support students. 

Despite the excitement and potential, there are several critical barriers to the success of AI tutoring. For one, out-of-the-box Language Models (LMs) often generate “bad pedagogical responses,” such as providing the answer too quickly, which can actually harm learning by robbing students of the chance to think for themselves (von Hippel, 2024). Tutors in the CoPilot study also flagged that AI suggestions were sometimes “too smart” or not grade-level appropriate, requiring the human tutor to simplify the language for the student (Wang, et al., 2024). Furthermore, LMs often lack real-world knowledge about specific curricula, as well as a student’s unique prior knowledge and experience, all of which are strengths that human educators bring to the table. And while Stanford’s AI-supported tutoring improved proximal mastery (like exit tickets), the two-month study did not find statistically significant improvements in end-of-year state math test scores (Wang et al., 2024).

Perhaps most importantly, the push for tailored learning illustrates several broader misconceptions about learning that are the backbone of these movements. For one, tailored learning often assumes that students possess fundamentally different “learning styles” or “paces” that require independent pathways (Bell, 2026). Yet decades of cognitive science (Kirschner et al., 2006) and classroom research (Rosenshine, 2012) show that humans share a common learning architecture: we all rely on limited working memory, build knowledge in long-term memory, and learn through explicit instruction, practice and feedback. Students may differ in background knowledge and opportunity, but those differences affect how quickly they learn, not how they learn. Furthermore, Benjamin Bloom’s vision for mastery learning was actually designed to eliminate initial differences in learning speed (Kulik et al., 1990). He predicted that with proper feedback, the correlation between a student’s initial aptitude and their final achievement would drop to near zero, allowing all students to eventually learn at the same quick pace (Kulik et al., 1990). Yet a TNTP (2018) study found that when students of all backgrounds were given the opportunity to work on grade-appropriate assignments, with the support of a rigorous teacher, they were capable of meeting those standards more often than not. While advocates argue for “personalized pacing,” meta-analyses show that group-based models (where students move together) actually produced larger positive effects(g=0.58) than self-paced programs (g=0.48) (Kulik et al., 1990).

Another misconception that perpetuates tailored learning is the belief that a single teacher cannot manage students at multiple readiness levels (Bell, 2026). Yet research consistently shows that preparation gaps are most effectively closed through coherent, high-quality instruction rather than isolation (TNTP, 2018). The Tutor CoPilot study further suggests that the problem isn’t a coherent classroom issue, but access to expert pedagogical reasoning (Wang et al., 2024). Even Bloom’s illustrious two-sigma study consisted of a “cocktail” of interventions, including extra instructional time, constant corrective feedback, and retesting that whole-class students did not receive (von Hippel, 2024). 

If nothing else, the past forty-years of iterations of tailored learning have made one thing abundantly clear: equity is achieved through high expectations and instructional coherence, not through technology-mediated independence that leaves the most vulnerable students to self-regulate their own remediation. The biggest issue facing our students today is not a lack of innovative ideas, but a lack of high-quality, well-prepared, and fully supported teachers, who are the most critical resource for student success. In an educational landscape that often undermines teacher expertise, fails to prepare them for classroom realities, and then adds a crushing implementation burden through fragmented “tailored” models, it is no wonder that the teacher pool is rapidly shrinking (Learning Policy Institute, 2025). Rather than paying millions of dollars for AI and tailored ed-tech tools, we must focus on equipping educators with evidence-based strategies while they are still in teacher preparation programs (Peske, 2026). This lack of foundational training makes the complex task of scaffolding instruction for students at different levels nearly impossible to master, leading schools to default to personalized dashboards that simply lower expectations. 

When schools adopt these tailored models to compensate for instructional gaps, the burden of managing these complex systems falls heavily on the faculty. In these models, teachers must move away from coherent, whole-class instruction to manage 30 independent pathways, often acting as administrators of digital dashboards rather than instructional leaders. Perhaps most disturbing is that a significant tension exists here between viewing a teacher as a reflective professional and viewing them as a technical functionary within a “human engineering” system. This is particularly concerning when the impact of teacher expectations is one of the most influential factors in student growth (TNTP, 2018). A tech-driven system fails to send teachers the message that mindset matters nearly as much as the material they teach. 

One has to wonder what Winston Smith would have thought about the robot, Plato: a personified educator that is “always patient, always available, and capable of adapting in real-time to a student’s pace and emotional state” (The White House, 2026). Of course, this also rests on faulty assumptions about what effective instruction looks like. More broadly, the danger in education is not that technology will become Big Brother. It is that systems can become so fluent in the language of innovation that they make it harder for people to ask independent, uncomfortable questions about whether students are actually learning; and whether the allure of big tech really is all it’s cracked up to be. The reality is that “meeting students where they are” has resulted in widening achievement gaps, particularly for low-income families and students of color (TNTP, 2018). We must focus our efforts on preparing teachers appropriately for the realities of the classroom, with evidence-based instructional practices because equity is first and foremost, a HUMAN ENDEAVOR. Student success depends on having a teacher who believes they can meet a high bar and has the skills and knowledge to help them get there. Investing in tech at the expense of teacher skill is, a cautionary tale where ignorance is strength. 

References

Alvarez, J. I., & Angeles, J. R. (2025). Khanmigo in the virtual classroom: A strategic evaluation through SWOT and acceptability analysis. Educational Process: International Journal, 16, Article e2025272. https://doi.org/10.22521/edupij.2025.16.272

Barnett, B. (2026, March 16). Khan Academy’s Khanmigo after one year: What the data actually shows about AI tutoring in schools. Edrus. https://edrus.org/

Bell, L. (2026, March 26). State leaders explore an education system based on student competency instead of seat time. EdNC. https://www.ednc.org/3-26-2026-state-leaders-explore-an-education-system-based-on-student-competency-instead-of-seat-time/

Cheng, L., Ritzhaupt, A. D., & Antonenko, P. (2019). Effects of the flipped classroom instructional strategy on students’ learning outcomes: A meta-analysis. Educational Technology Research and Development, 67, 793–824. https://doi.org/10.1007/s11423-018-9633-7

Evans, C. M., Landl, E., & Thompson, J. (2020). Making sense of K-12 competency-based education: A systematic literature review of implementation and outcomes research from 2000 to 2019. Journal of Competency-Based Education, 5(4), e01228. https://doi.org/10.1002/cbe2.1228

Kirschner, P. A., Sweller, J., & Clark, R. E. (2006). Why Minimal Guidance During Instruction Does Not Work: An Analysis of the Failure of Constructivist, Discovery, Problem-Based, Experiential, and Inquiry-Based Teaching. Educational Psychologist41(2), 75–86. https://doi.org/10.1207/s15326985ep4102_1

Kulik, C. C., Kulik, J. A., & Bangert-Drowns, R. L. (1990). Effectiveness of mastery learning programs: A meta-analysis. Review of Educational Research, 60(2), 265–299. https://doi.org/10.3102/00346543060002265

Learning Policy Institute. (2025, July 16). An overview of teacher shortages: 2025 [Fact sheet]. Learning Policy Institute.https://learningpolicyinstitute.org/product/overview-teacher-shortages-2025-factsheet

Meyer, D. (2026, April 15). RIP Khanmigo & edtech industry dreams of AI tutors. Amplify. https://www.linkedin.com/pulse/rip-khanmigo-edtech-industry-dreams-ai-tutors-dan-meyer-bfuec/

Ramírez Naranjo, N. (2022). Criticisms of the competency-based education (CBE) approach. In A. Opačić (Ed.), Social work in the frame of a professional competencies approach (pp. 21–37). Springer Nature Switzerland AG. https://doi.org/10.1007/978-3-031-13528-6

Pane, J. F., Steiner, E. D., Baird, M. D., Hamilton, L. S., & Joseph, D. P. (2017). How does personalized learning affect student achievement? (Research Brief). RAND Corporation. https://www.rand.org/pubs/research_briefs/RB9994.html

Peske, H. (2026, April 23). Holding teacher prep programs accountable is hard. States should do it anyway. National Council on Teacher Quality. https://www.nctq.org/research-insights/holding-teacher-prep-programs-accountable-is-hard-states-should-do-it-anyway/

Rosenshine, B. (2012). Principles of instruction: Research‑based strategies that all teachers should know [PDF]. American Educator, 36(1). https://www.aft.org/sites/default/files/Rosenshine.pdf

The White House. (2026, March 25). First lady Melania Trump convenes record 45 nations at the White House and introduces American-built humanoidhttps://www.whitehouse.gov/briefings-statements/2026/03/first-lady-melania-trump-convenes-record-45-nations-at-the-white-house-and-introduces-american-built-humanoid/

TNTP. (2018). The opportunity myth: What students can show us about how school is letting them down—and how to fix ithttps://tntp.org/publication/the-opportunity-myth/

von Hippel, P. T. (2024). Two-sigma tutoring: Separating science fiction from science fact. Education Next, 24(2), 22–31.https://www.educationnext.org/two-sigma-tutoring-separating-science-fiction-from-science-fact/

Wang, R., Zhang, Q., Robinson, C. D., Loeb, S., & Demszky, D. (2024). Scaling expert-like guidance for novice tutors(ArXiv Preprint 2410.03017v2). https://doi.org/10.48550/arXiv.2410.03017


Leave a Reply

Discover more from Rebecca A. Huggins

Subscribe now to keep reading and get access to the full archive.

Continue reading