One question organizes this research program: what cultures, contexts, and learning experiences give people the knowledge, skills, beliefs, and motivation to become effective users of Generative AI (GenAI), tools like ChatGPT — and do so equitably?
Effective use is not an amount and not a trait. The same student works with AI to think in one class and to avoid thinking in another; the same amount of use serves opposite goals. It is a fit between what a person does with the tool, the goal they hold, and what the situation asks of them — which means it has to be measured as a fit and taught as a judgment.
Answering the question takes five things the field does not yet have:
- a shared understanding of what helpful and harmful use actually are;
- measures that capture the quality of use — how, not how much;
- research methods that elicit authentic variability in type of use, rather than averaging it away;
- evaluation infrastructure built alongside deployment instead of after it;
- capacity in the field and the public to conduct and judge the science.
Then the answers have to become programs and policies that are evaluated, revised, scaled, and kept under evaluation as they scale.
The clock. GenAI-and-learning research is producing stories faster than the field can absorb them — thousands of papers and dozens of meta-analyses since 2023 — a miracle tutor in one telling, cognitive decay in the other. Both are hardening into policy while the evidence base is still forming. The field has run this exact sequence before — twice.
The field has been here before
The cost of scaling too fast. In the 1990s, self-esteem was promoted as a social vaccine: raise children's self-esteem and failure, delinquency, and drug use would fall away. It is a familiar story: excitement over a few findings that were not designed to support the conclusion, combined with skipping the science needed to build, test, and scale the pipeline from promising finding to program. The promise organized programs, curricula, and public money for a generation, while the backlash that followed stunted progress on understanding how self-esteem should be conceived and measured, how to foster healthy levels of it, for whom, and under what conditions. I began my career inside that backlash — studying self-esteem, and defending the worthiness of the work at every step.
The power of a good story versus data. The story then evolved: the self-esteem movement had supposedly turned a generation of students into “Generation Me” — more confident, more assertive, more entitled, and more miserable than any before it (Twenge, 2006) — and then into a full epidemic of narcissism (Twenge & Campbell, 2009). All of it built on conclusions that spoke beyond the data, and all of it forming stories the world felt made sense. Young people absorbed the stereotype and applied it to themselves (Trzesniewski & Donnellan, 2014). The press argued that the narrative must be true while sparing the data a casual glance; the public argued from samples of one. And the stories persisted and grew against an academic debate that rarely changed established views: a five-year exchange of paper and counter-paper that left the popular narrative entirely intact (Twenge, Konrath, Foster, Campbell, & Bushman, 2008; Trzesniewski, Donnellan, & Robins, 2008; Twenge & Foster, 2008; Donnellan & Trzesniewski, 2009; Trzesniewski & Donnellan, 2010; Twenge & Campbell, 2010; Twenge, 2013; Arnett, Trzesniewski, & Donnellan, 2013).
Social media: the lessons to learn. Social media came next, invoked as the cause of not just one troubled generation but every one after it — another wave of programs, curricula, and public money invested in policy without the scientific foundation. Catalysts of that failure include: measurement innovation that lagged — screen time counted when what mattered was what was consumed, by whom, and from whom — leaving years of large longitudinal studies using assessments that provide limited insight; a focus on average effects, obscuring the individual differences and situations that determine whether a given use helps or harms; and little focus on identifying mechanisms in authentic use situations that a program could target. The result is a large literature that cannot answer the question policymakers are asking.
A working hypothesis, offered as one. I believe in the potential of GenAI to improve lives and reduce inequities — to put knowledge and mentoring in the hands of people usually left out when resources are distributed. That belief is the hypothesis this program exists to test, not a conclusion it starts from. The field has seen what happens when stories, programs, and policies get ahead of the knowledge — at the cost of a generation's self-understanding the first time and a decade of blunt policy the second. The third time is now, and it is moving faster.
Where this comes from, and where it is headed
Two strands, and why this is the merge. Two strands run through my colleagues' careers and mine. One develops ways for people to succeed in contexts that don't serve everyone — the person-level work: self-esteem, growth mindset, belonging, and the measures that make them studiable. It works, and it puts the burden in the wrong place: on students, to adapt and stay resilient in environments not built for them. And the benefit of even a well-run mindset intervention depends on whether the classroom culture lets a student enact it. So the other strand works to change the culture itself — the shared beliefs, norms, and goals that define what it means to be a learner in a classroom — giving students the affordances to enact what they bring, and lifting the burden of resilience off them alone. Each strand has taught its own lesson: the changes to classroom culture that matter most demand work instructors are rarely given the time for, and the person-level promise of GenAI is stalling because decades of learning science are not being brought into the work. This program is the merge: a conceptual model and measures, both in development on these pages, bring the two strands to the same question.
The program
Three lines of work, and they run at once: correct the record; develop the measures; build the environments, then scale what works. Correcting the record shows which designs cannot support the claims that traveled, and each recurring failure names a requirement the measures and designs have to meet. Building real learning environments is what makes the full range of use appear — the only condition under which those measures can be validated. The research standards that correcting the record feeds cannot be written before the science exists, because what counts as a valid measure, or an authentic learning environment, is itself something the science has to discover. And writing the standards and training reviewers to use them is where the field's capacity to conduct and judge this science gets built.
I am beginning to document all three lines here — what the evidence supports, what it does not yet, and what has already had to be revised. If you are thinking about similar things, please reach out (ktrz@ucdavis.edu) — I would love to compare notes, debate, and brainstorm.
Go deeper
-
01
Correct the record
The most talked-about studies cannot support their headlines; the counter-evidence is not being read. Where the record actually stands, and how it gets fixed upstream.
-
02
Develop the measures
Define effective use, build the framework, develop the measures: the questions the field had not put together, what we are learning, and the validation road ahead.
-
03
Build the environments, then scale what works
Effective use has to be learned somewhere, and measured somewhere. What stops those settings from being built — the instructor's labor, and evidence that never reaches them — and how GenAI is being used to lower both barriers.
References
- Arnett, J. J., Trzesniewski, K. H., & Donnellan, M. B. (2013). The dangers of generational myth-making: Rejoinder to Twenge. Emerging Adulthood, 1, 17–20.
- Donnellan, M. B., & Trzesniewski, K. H. (2009). An emerging epidemic of narcissism or much ado about nothing? Journal of Research in Personality, 43, 498–501.
- Trzesniewski, K. H., & Donnellan, M. B. (2010). Rethinking “Generation Me”: A study of cohort effects from 1976–2006. Perspectives on Psychological Science, 5, 58–75.
- Trzesniewski, K. H., & Donnellan, M. B. (2014). “Young people these days…”: Evidence for negative perceptions of emerging adults. Emerging Adulthood, 2, 211–226.
- Trzesniewski, K. H., Donnellan, M. B., & Robins, R. W. (2008). Is “Generation Me” really more narcissistic than previous generations? Journal of Personality, 76, 903–918.
- Twenge, J. M. (2006). Generation Me. Free Press.
- Twenge, J. M. (2013). The evidence for Generation Me and against Generation We. Emerging Adulthood, 1, 11–16.
- Twenge, J. M., & Campbell, W. K. (2009). The Narcissism Epidemic: Living in the Age of Entitlement. Free Press.
- Twenge, J. M., & Campbell, W. K. (2010). Birth cohort differences in the Monitoring the Future dataset and elsewhere: Further evidence for Generation Me. Perspectives on Psychological Science, 5, 81–88.
- Twenge, J. M., & Foster, J. D. (2008). Mapping the scale of the narcissism epidemic: Increases in narcissism 2002–2007 within ethnic groups. Journal of Research in Personality, 42, 1619–1622.
- Twenge, J. M., Konrath, S., Foster, J. D., Campbell, W. K., & Bushman, B. J. (2008). Egos inflating over time: A cross-temporal meta-analysis of the Narcissistic Personality Inventory. Journal of Personality, 76, 875–902.