Ingestion debt
The most addicted AI users I know share the same tell. They send the next prompt before the last answer has even finished generating. Not before they finished reading it, before the model finished writing it.
The answer was good. It usually is now. It streamed to its end in a viewport nobody was watching, while the follow-up was already being typed. Multiply that by a workday, then by a year. Chat logs nobody will reopen. Summaries saved to read later, where later is a place things go to be forgotten. Forty tabs, each one a small promise.
Every one of those unread answers is a loan. You borrowed the feeling of knowing and deferred the work of actually knowing, and the balance compounds quietly, in re-asked questions, re-derived conclusions, decisions made on text nobody absorbed. Call it what it is, ingestion debt.
Generation became free. The price of understanding inflated. Understanding is paid in attention, and attention got scarcer while the text demanding it went infinite. That asymmetry is the defining constraint of working and learning in 2026, and almost every tool in your stack is built to widen it.
I spent the last month reading the actual evidence on how humans durably absorb what they read, 44 primary sources, meta-analyses and randomized trials and official data, every claim checked against the original paper. One viral meta-analysis died of retraction while I was reading it. This essay is the whole thing, the debt, the science, the traps, and the repayment loop, in one place.
A note on reading it, because the topic demands it. This is a long essay, about twenty minutes. The bolded lines are chosen deliberately, provided highlights outperform your own (Ponce 2022), so skimming them is a legitimate first pass. The sections are anchored if you need to leave and come back.
TipGuess first. Three questions.
Before you read on, commit to an answer for each. Wrong guesses are not a problem, guessing first measurably improves what you retain from the text that follows. That claim has a citation, and you will meet it in section four.
- Which single study habit does the evidence rank first for long-term retention, and do you actually use it?
- Does listening to a book retain worse than reading it?
- Does highlighting help you remember what you highlighted?
Your answers get checked at the end.
The debt#
Attention on any single screen, before switching to another one, averaged two and a half minutes in 2004. By the early 2020s it was 47 seconds, median 40. That number predates the current AI wave. It describes the reader that AI-generated text now arrives to, in unprecedented volume, at zero marginal cost.
What happens when that reader writes with a model instead of ingesting? An MIT Media Lab team took EEG measurements of people writing essays with and without an LLM. In the assisted group, 83 percent could not quote a single sentence from the essay they had just written (Kosmyna et al. 2025). Minutes after finishing it. The deficit persisted when they were asked to work unassisted afterward. It is a preprint with a small sample, so treat it as directional, but the direction is the one every heavy AI user will recognize from the inside. The output existed. The encoding never happened.
The cleanest causal evidence comes from a randomized trial in high school math, published in PNAS (Bastani et al. 2025). Students with unrestricted GPT-4 access solved 48 percent more practice problems correctly. On the closed-book exam that followed, they scored 17 percent worse than students who never had the tool. A guardrailed tutor version, built to coach rather than answer, inflated practice performance by 127 percent and managed only to bring the exam harm back to zero. Read that carefully, because it is the whole problem in one design. Assistance inflated the appearance of learning while the learning itself went negative, and careful guardrails only broke even.
The mechanism is not mysterious. Producing an answer yourself is the thing that writes it into memory. When the machine produces it for you, the effort disappears, and the encoding disappears with it. I keep meeting the same mechanism in code, where engineers merge agent-written changes they never really read and pay for it weeks later, a comprehension debt with the same interest schedule. My close friend Mert Cobanov has been making this point about systems for as long as I have known him. Every migration promised for later is a loan, and the longer later takes to arrive, the more has quietly compounded by the time you finally pay. It is one asymmetry wearing many costumes. The model
was never the bottleneck. The absorbing layer is.
None of this is an argument against AI, and this essay is not a nostalgia piece. It is an argument about which side of the pipe is now scarce. Text got infinite. Encoded understanding did not. Every unread answer is a loan, and the interest is paid in re-work.
Exposure is not encoding#
The debt metaphor only holds if exposure and encoding really do come apart. They do, and the literature on it is old, replicated, and brutal. Andy Matuschak made the argument famous in Why books don't work, the book as a medium quietly assumes that people absorb what they read, and almost nobody does. What follows is the measured version of that claim.
Start with what your eyes are doing right now. In studies of mindless reading, the eyes keep moving across the page in a perfectly normal pattern while attention has decoupled entirely (Smallwood 2011). Gaze position says reading. The mind is elsewhere. Across studies, readers spend 20 to 30 percent of reading time mind-wandering, and the more they wander, the less they comprehend (
D'Mello & Mills 2021), a correlation around minus 0.3. A fifth to a third of the time you spend reading is exposure without any encoding at all.
Worse, you cannot feel it happening. When learners grade their own recall, the judgment "I got that right" is correct only 57 percent of the time (Dunlosky & Rawson 2012). In the same research program, students who over-trusted that feeling retained 32 percent of the material, and students trained to check themselves against an objective standard retained 89 percent. Same material, same time invested. The difference was calibration. The feeling of understanding is a lagging, lying indicator.
This is why rereading, the world's default study move, keeps losing experiments. In the classic head-to-head, students who reread a passage predicted they would remember it best. A week later the group that had instead spent that time being tested on the passage recalled 56 percent against the rereaders' 42 (Roediger & Karpicke 2006). The rereaders felt better and performed worse, and that inversion is the single most important fact in this entire literature.
The positive version of the fact is called retrieval practice, and it is about as close to settled as this field gets. Three meta-analyses from three different literatures converge. Across 61 lab studies, testing beat restudying at g = 0.50, and the advantage grew as the retention interval got longer (Rowland 2014). Across 217 studies,
g = 0.61. Across 48,478 real students in real classrooms, quizzing lifted achievement by half a standard deviation (
Yang et al. 2021). Two moderators matter enough to state as rules. Feedback roughly doubles the effect, and without feedback, testing at low success rates does
nothing at all, g = 0.03. Retrieval is the spine. Feedback is the vertebrae.
So the mechanics are not in dispute. Memory is built when you produce the material, under some difficulty, with a correction available, and it decays when you merely re-expose yourself to it, however pleasant the re-exposure feels.
NoteCheckpoint
Look away from the screen and say, in one sentence, what this section claimed. If nothing comes back, that is not a failure, it is the section's point demonstrating itself. The claim was, the feeling of understanding and the fact of it are different signals, and only retrieval tells you which one you have.
The graveyard of reading advice#
Before building on the part of the evidence that holds, clear out the part that does not. Reading and learning advice is unusually thick with confident claims that died in replication, and knowing which ones is half the value of doing the research at all.
| The claim | What the evidence says | Source |
|---|---|---|
| Match teaching to your learning style | No credible evidence, after an adversarial review designed to find some | |
| Brain training makes you smarter | Gains stay on the trained task. The FTC fined Lumosity over its claims | |
| Working-memory training transfers | Far transfer is d ≈ 0.00 with proper controls | |
| Speed reading works | A speed-accuracy trade. RSVP kills the re-reading that repairs comprehension | |
| Handwriting beats typing for notes | Failed replication, pooled g = 0.04. What predicts recall is content captured, not the medium | |
| Highlighting is studying | Your own highlights aid memory a little, g ≈ .36, and comprehension not at all | |
| Expanding review intervals beat fixed ones | Folklore. g = 0.034, not significant. Spacing itself is what matters | |
| Micro-breaks restore attention | The famous brief-breaks result failed direct replications | |
| "ChatGPT boosts learning, g = 0.867" | The viral meta-analysis behind it was retracted in April 2026 | |
| "Body doubling works for 85% of ADHDers" | The statistic is fabricated. The practice is plausible but has zero controlled trials |
That last pair deserves a beat. A meta-analysis claiming enormous learning gains from ChatGPT went viral, got cited everywhere, and was withdrawn. A body-doubling statistic circulates in every ADHD community and traces back to nothing. Most reading advice is folklore with a citation. The only defense is reading the paper, which is why every number in this essay links to one.
What survived#
Strip away the graveyard and the surviving techniques share one property. Every one of them moves the learner from receiving material to producing it. That is the through-line, and it has a name in the literature, the generation effect, d = .40 across 86 studies, rising to .64 once a day has passed. The rest is engineering around that one fact.
| Law | The evidence | In practice |
|---|---|---|
| Retrieve, do not re-expose | Three metas, g = 0.50 to 0.61, | Close the tab. Say what it claimed. Then check |
| Space by the horizon | Optimal gap is 10 to 20 percent of how long you need it, about 5 percent at a year ( | Review tomorrow, then next week, then next month |
| Feedback after every failure | With feedback g = 0.73, without 0.39, and failed retrieval without correction is | Never end on a wrong answer |
| Checkpoint through, not after | Interpolated tests cut mind-wandering from 39 to 19 percent and lifted final recall ( | A question every ten minutes beats a quiz at the end |
| Guess before you read | Pretesting beats post-testing, d = 0.30, when reading follows the guess ( | The three questions at the top of this essay |
| Ask exactly what you want to keep | Tutoring gains collapse from 0.73 to 0.13 when the test does not match the practice ( | Write the question you actually need answered later |
| Grade yourself objectively | Self-graded "correct" is right 57 percent of the time ( | Type the answer, then compare, no vibes |
| Self-pace, whatever the modality | Reading and listening are equivalent when you control the pace ( | Audio is fine. Skimming against a clock is not |
| Schedule with a learned model | Learned schedulers beat the 1987 algorithm for 99.6 percent of ten thousand users (SRS Benchmark). Personalized spacing beat generic by a letter grade ( | Use FSRS, not SM-2, not intuition |
| Forgive the streak | Habits take 66 days at the median, range 18 to 254, and single missed days cost nothing ( | A missed day is noise. Punishing it is how systems die |
Two of these deserve expansion because they run against instinct.
First, modality does not matter nearly as much as control over pacing. The reading-versus-listening meta found no overall difference, and the small edge that appears for text shows up mainly where readers can regress and pause and listeners cannot. This matters for anyone who reaches for audio or text-to-speech because their eyes slide off dense pages. The evidence says that is a legitimate accommodation, not a lesser way to read, provided the pause button gets used like a tool.
Second, the checkpoint law is the direct answer to the 20-to-30-percent wandering tax. In the interpolated-testing studies, inserting brief retrieval between segments halved mind-wandering and roughly doubled final recall (Szpunar et al. 2013), where merely re-presenting the material did nothing. Structure beats willpower. You do not fix drift by trying harder, you fix it by making the text ask things back.
Produce the memory. Do not receive it. Every law above is that sentence wearing different clothes.
NoteCheckpoint
Same drill as before. Name two laws from the table without scrolling up. Notice how different this feels from the smooth confidence of reading them, that gap is the fluency illusion from section two, experienced live.
You prefer what fails you#
Here is the finding that explains why, if all of this is so well-established, nobody does it and no mainstream product ships it.
In the spacing literature, learners who experienced both schedules and were then shown their own results still got it backwards. 72 percent judged cramming to have worked better even after spacing had objectively improved 90 percent of them (Bjork, Dunlosky & Kornell 2013). The judgment tracks how learning feels while it happens, and effective learning feels worse. Effortful, slow, full of failed retrievals. The field calls these
desirable difficulties, and the point of the name is that learners reliably avoid them.
The AI version of this experiment already exists. A 2025 study had people study with AI note-taking at three automation levels. Fully automated notes produced the lowest test scores and the highest preference ratings (2025 preprint). Participants rated the condition that taught them least as best on quality, enjoyment, and intent to use. The best-scoring condition, where AI proposed atomic notes and the human had to assemble and edit them, was liked least. Small study, 30 people, but it lands exactly on the 72-percent result from a completely different literature. Preference and learning are not just uncorrelated. In this domain they point in opposite directions.
Now look at the product landscape with that lens. Summary apps sell compression, read less, know more, which the exam data says is exactly backwards, and the summary category hit its ceiling years ago while subscription learning kept growing around it. Auto-flashcard tools sell the removal of the one effortful step that made the cards work. Beautiful canvas tools sell organization to the already-motivated, a profitable niche that never touches the person whose problem is attention itself. And the tools that do center real retrieval, Anki and its lineage, present like homework software and are abandoned accordingly. The market keeps building vending machines for the feeling of learning because that is what the buyer's thumb votes for.
The one mainstream product that cracked it is instructive. Duolingo runs streaks, leagues, and a learned spacing model over what is, underneath, plain spaced retrieval, and reports a 41 percent daily-to-monthly active ratio (Q1 FY26 letter), retention that consumer apps do not get. Its own research team also published the honest caveat, the famous engagement gains came from better streak mechanics, and the memory model's job was cutting prediction error, not producing the viral number (
Settles & Meeder 2016). The lesson is not gamify everything. A meta-analysis of gamification found the cognitive benefit
survives rigorous designs while bare points-and-badges effects wash out. The lesson is that adherence machinery around a real retrieval engine is the only combination with receipts.
The market sells the feeling. The learning lives in the friction. Any honest system, personal or product, has to resolve that tension rather than pick a side of it.
The loop#
Everything above compresses into one operating loop. This is the repayment plan for ingestion debt, and it is deliberately boring, six steps, each one carrying a specific finding on its back.
1. Capture into one inbox. Papers into a reference manager, articles into one read-later lane, everything with its source attached. Capturing is cheap and fine. The rule that makes it safe is that capture reserves the option to read, it is not a commitment to read. The collector's fallacy starts where that line blurs.
2. Triage without mercy. For each item, one verdict. Ignore, skim, read, mine for one specific claim, or schedule deep work. The active reading queue stays at five to nine items. Beyond that a queue stops being a plan and becomes guilt storage, and guilt storage is where reading systems go to die.
3. Read with scaffolds, not willpower. Write the purpose before starting, one line, what decision or project does this feed. Guess the answers first, the pretesting effect pays out even when the guesses are wrong. Then checkpoint every ten minutes or every section, look away, state the claim, note what would change your mind. That single habit attacks the
mind-wandering tax at its measured source. Dense text through ears instead of eyes is fine, self-pacing is the ingredient that matters (
Clinton-Lisell 2022).
4. Convert claims into questions. End every serious read by writing one to five questions, not summaries, questions with checkable answers. This is where AI belongs in the loop, drafting candidate questions from what you highlighted, detecting contradictions against other sources, asking Socratic follow-ups. The one job it must never take is answering the questions for you, because the assembly is the encoding (
2025 preprint). Hard cap the volume. Five cards per article, twenty per book chapter. A review queue that outgrows its reader is just the debt again in a new costume.
5. Review on a learned schedule. The questions go into spaced retrieval with FSRS doing the scheduling. Type or say the answer before revealing, then grade against the shown answer, which repairs the 57-percent self-grading problem. Cap daily review time and cut low-value cards without sentiment. Missed a day,
nothing happened, continue.
6. Link what survives into notes you can find. One idea per note, source attached, connected to the project it serves. This layer is what makes the whole loop compound, and it only compounds if every writer, human or agent, follows the same shape.
The loop's power-user version has a name in the literature, successive relearning, retrieve to criterion in one session, then again days later, then again. In real courses with real exams it produced gains of half to a full letter grade, d = 0.54 to 1.10 (Janes et al. 2020), one of the largest effects ever measured in situ. The same paper carries the warning that shapes everything I am about to say next. When the practice was optional, 61 percent of students used it zero times. The technique was never the hard part. Capture less. Retrieve more. The hard part is architecture that makes you show up.
What I am building#
I did not run this research pass out of general curiosity. I went in with a product intuition, an infinite canvas of two-minute linked cards for readers whose focus fractures, and I expected the evidence to bless it. It did something more useful. It killed the flattering parts and left the load-bearing ones.
The canvas, it turns out, is commodity. The retention mechanics are the moat. The evidence-mandated shape is a reading companion where the loop itself is the product. It ingests whatever you actually read, segments it around natural stopping points because attention budgets are measured in seconds now, opens each segment with a guess-first question, checkpoints through the text, drafts the retrieval questions for your confirmation, one tap to accept or edit, because
user-assembled beats auto-generated and auto-generated is the trap. Then it schedules everything on a learned model and grades recall objectively. The metric it optimizes, and the only one on its dashboard, is what you still know at seven days, because
the quiz-day number lies and time-on-page lies harder. The closest prior art is Matuschak's
mnemonic medium, essays with the retrieval woven directly into the reading, and the bet here is carrying that pattern from motivated quantum-computing students to the reader whose problem is attention itself.
The audience it serves first is the one every incumbent ignores, the focus-challenged reader. That population is not a niche. 6.76 percent of adults worldwide have symptomatic ADHD (Song et al. 2021), around 366 million people, closer to nine percent in the young-adult cohort, and prescription data shows demand surging, fastest among adult women (
CDC 2023). The digital tools marketed at them have, per the umbrella reviews,
inconclusive evidence at best. What the learning literature offers instead is scaffolding with receipts, checkpoints, pretesting, forgiving schedules, and that is buildable.
I hold the claims of this section more loosely than the science above, and the open questions are stated in the research, honestly, as A/B tests rather than beliefs. Whether the ideal segment is two minutes or ten. Whether a linear next-step chain helps or fights the curiosity that actually sustains attention. Whether effortful retrieval can be made to feel light enough for the audience that needs it most, which is the make-or-break design problem, because the preference trap does not care how good the evidence is.
Now settle the three guesses from the top. The habit the evidence ranks first is retrieval practice, testing yourself, and if you are like most readers you reread instead. Listening does not retain worse than reading, provided you control the pace. Highlighting barely helps memory and does not help comprehension, though reading someone's provided highlights does, which is what the bold lines in this essay were.
If your three answers were already right, you are ahead of the illusion this essay is about. If any surprised you, that gap between what you believed and what the evidence shows is ingestion debt in miniature, and you now know the repayment mechanics.
Generation is only getting cheaper. The models will keep improving, the answers will keep arriving faster than anyone reads them, and the spread between what passes through you and what stays will keep widening on its own. Retention is the ledger that matters now. Run the loop.
The shelf#
The 44 sources behind this essay, grouped by what they explain, one line each. Every claim above traces to one of these, and they are worth more than my compression of them.
A. Memory and retrieval, 9 sources
Rowland 2014, meta-analysis. Testing beats restudy, g = 0.50, growing with delay, worthless without feedback at low success.
Adesope, Trevisan & Sundararajan 2017, meta-analysis. g = 0.61 overall, multiple-choice is legitimate, classroom matches lab.
Yang et al. 2021, meta-analysis. Classroom quizzing g = 0.499 across 48,478 students.
Cepeda et al. 2008, N = 1,354. Optimal review gap is 10 to 20 percent of the retention horizon. Err long.
Latimier, Peyre & Ramus 2021, meta-analysis. Spaced retrieval beats massed, g = 1.01. Expanding intervals are folklore, g = 0.034.
Brunmair & Richter 2019, meta-analysis. Interleaving helps confusable categories, g = 0.42, not prose.
Janes et al. 2020, RCT. Successive relearning lifted real exams 13 percent, d = 0.54 to 1.10. Optional practice went 61 percent unused.
Pan & Sana 2021, five RCTs. Guess-first beats quiz-after, d = 0.30, when reading follows the guess.
Bertsch et al. 2007, meta-analysis. Generation effect d = .40, rising to .64 past a day.
B. Attention and mind-wandering, 6 sources
Szpunar, Khan & Schacter 2013. Interpolated tests halve mind-wandering, 19 versus 39 to 41 percent.
Gruber, Gelman & Ranganath 2014. Curiosity states lift recall, 70.6 versus 54.1 percent, and spill onto incidental material.
Ariga & Lleras 2011. The brief-breaks vigilance result, later failed direct replications.
Mark 2023, APA. On-screen attention before switching, 2.5 minutes in 2004 to about 47 seconds.
Smallwood 2011. Mindless reading, eyes scan while attention decouples.
D'Mello & Mills 2021. Mind-wandering fills 20 to 30 percent of reading, r ≈ −.3 with comprehension.
C. Modality, screens, audio, notes, 8 sources
Delgado et al. 2018, meta-analysis. Screen inferiority g = −.21, three times larger under time pressure, informational text only.
Rayner et al. 2016, review. Speed reading is a speed-accuracy trade. RSVP breaks comprehension repair.
Clinton-Lisell 2022, meta-analysis. Reading equals listening overall. Self-pacing is the active ingredient.
Wood et al. 2018, meta-analysis. Text-to-speech helps reading-disabled students, d = .35, corrected .24. An accessibility lever.
Ponce, Mayer & Méndez 2022, meta-analysis. Own highlights, memory .36, comprehension nothing. Provided highlights, .44 on both.
Morehead, Dunlosky & Rawson 2019, RCT. Pen versus keyboard failed to replicate.
Rogowsky et al. 2016, RCT. Audiobook equals e-text equals both, immediately and at two weeks.
Mueller & Oppenheimer 2014, with Urry et al. 2021. The medium myth died, pooled g = 0.04. Verbatim capture equals shallow processing survived.
D. Metacognition and the myths, 6 sources
Bjork, Dunlosky & Kornell 2013, review. 72 percent judge massing better even after spacing improved 90 percent of them.
Pashler et al. 2008, adversarial review. Learning-styles matching has no credible evidence.
Melby-Lervåg & Hulme 2013, meta-analysis. Working-memory training, far transfer d ≈ 0.00.
Bjork & Bjork 2011. Desirable difficulties, with the boundary that difficulty must match prior knowledge.
Dunlosky & Rawson 2012. Self-graded correct is right 57 percent of the time. Calibration moves retention 32 to 89 percent.
Simons et al. 2016, review. Brain training improves the trained task, little else. The FTC fined Lumosity.
E. AI and algorithmic learning, 7 sources
Bastani et al. 2025, PNAS, RCT. Unguardrailed GPT help, practice +48 percent, exam −17. Guardrails neutralize, do not add.
Wang & Fan 2025, meta-analysis, retracted April 2026. The viral ChatGPT g = 0.867 number is withdrawn.
Kulik & Fletcher 2016, meta-analysis. Intelligent tutoring g ≈ 0.50 to 0.66, collapsing to 0.13 on unaligned tests.
Kosmyna et al. 2025, preprint. 83 percent of LLM-assisted writers could not quote their own essay. Directional.
Lindsey et al. 2014, classroom RCT. Personalized spacing beat generic by 8 to 10 percent, d ≈ 0.9 to 1.4. Massed won the quiz and lost the exam.
Settles & Meeder 2016, Duolingo. Trainable forgetting curves cut prediction error 45 percent. The famous engagement gain came from streak mechanics.
- SRS Benchmark, open dataset. Learned schedulers beat SM-2 for 99.6 percent of about ten thousand users, roughly 350 million reviews.
F. ADHD and the market, 8 sources
Lally et al. 2010, longitudinal. Habits form in about 66 days, range 18 to 254. Single missed days are harmless.
Gabarron et al. 2025, umbrella review. Digital ADHD interventions, inconclusive due to low-quality evidence.
Kollins et al. 2020, RCT. EndeavorRx moved its trained attention metric, none of the clinical secondaries.
Song et al. 2021, meta-analysis. Adult symptomatic ADHD, 6.76 percent, about 366 million adults.
CDC MMWR 2023. Stimulant fills surged 2020 to 2021, fastest among young adult women.
Sailer & Homner 2020, meta-analysis. Gamification's cognitive effect survives rigorous designs, bare points-and-badges effects do not.
Eagle et al. 2024, survey. Body doubling, widely practiced, zero controlled trials, the viral statistic fabricated.
Duolingo Q1 FY2026 shareholder letter. The habit loop compounds, 41 percent DAU over MAU.
The research itself, five documents and all 44 source summaries with limitations and replication notes, lives in my vault, and this essay is its step six, the synthesis that gets published. If you read all of this and remember one sentence a week from now, let it be the one you produced yourself at a checkpoint. That is the whole theory, demonstrated.