·

Ingestion debt

LearningAIMemoryReadingAdhdResearch
Contour artwork for Ingestion Debt, an evidence review of AI-assisted learning

The most addicted AI users I know share the same tell. They send the next prompt before the last answer has even finished generating. Not before they finished reading it, before the model finished writing it.

The answer was good. It usually is now. It streamed to its end in a viewport nobody was watching, while the follow-up was already being typed. Multiply that by a workday, then by a year. Chat logs nobody will reopen. Summaries saved to read later, where later is a place things go to be forgotten. Forty tabs, each one a small promise.

Every one of those unread answers is a loan. You borrowed the feeling of knowing and deferred the work of actually knowing, and the balance compounds quietly, in re-asked questions, re-derived conclusions, decisions made on text nobody absorbed. Call it what it is, ingestion debt.

Generation became free. The price of understanding inflated. Understanding is paid in attention, and attention got scarcer while the text demanding it went infinite. That asymmetry is the defining constraint of working and learning in 2026, and almost every tool in your stack is built to widen it.

I spent the last month reading the actual evidence on how humans durably absorb what they read, 44 primary sources, meta-analyses and randomized trials and official data, every claim checked against the original paper. One viral meta-analysis died of retraction while I was reading it. This essay is the whole thing, the debt, the science, the traps, and the repayment loop, in one place.

A note on reading it, because the topic demands it. This is a long essay, about twenty minutes. The bolded lines are chosen deliberately, provided highlights outperform your own (Ponce 2022), so skimming them is a legitimate first pass. The sections are anchored if you need to leave and come back.

TipGuess first. Three questions.

Before you read on, commit to an answer for each. Wrong guesses are not a problem, guessing first measurably improves what you retain from the text that follows. That claim has a citation, and you will meet it in section four.

  1. Which single study habit does the evidence rank first for long-term retention, and do you actually use it?
  2. Does listening to a book retain worse than reading it?
  3. Does highlighting help you remember what you highlighted?

Your answers get checked at the end.

The debt#

Attention on any single screen, before switching to another one, averaged two and a half minutes in 2004. By the early 2020s it was 47 seconds, median 40. That number predates the current AI wave. It describes the reader that AI-generated text now arrives to, in unprecedented volume, at zero marginal cost.

What happens when that reader writes with a model instead of ingesting? An MIT Media Lab team took EEG measurements of people writing essays with and without an LLM. In the assisted group, 83 percent could not quote a single sentence from the essay they had just written (Kosmyna et al. 2025). Minutes after finishing it. The deficit persisted when they were asked to work unassisted afterward. It is a preprint with a small sample, so treat it as directional, but the direction is the one every heavy AI user will recognize from the inside. The output existed. The encoding never happened.

The cleanest causal evidence comes from a randomized trial in high school math, published in PNAS (Bastani et al. 2025). Students with unrestricted GPT-4 access solved 48 percent more practice problems correctly. On the closed-book exam that followed, they scored 17 percent worse than students who never had the tool. A guardrailed tutor version, built to coach rather than answer, inflated practice performance by 127 percent and managed only to bring the exam harm back to zero. Read that carefully, because it is the whole problem in one design. Assistance inflated the appearance of learning while the learning itself went negative, and careful guardrails only broke even.

The mechanism is not mysterious. Producing an answer yourself is the thing that writes it into memory. When the machine produces it for you, the effort disappears, and the encoding disappears with it. I keep meeting the same mechanism in code, where engineers merge agent-written changes they never really read and pay for it weeks later, a comprehension debt with the same interest schedule. My close friend Mert Cobanov has been making this point about systems for as long as I have known him. Every migration promised for later is a loan, and the longer later takes to arrive, the more has quietly compounded by the time you finally pay. It is one asymmetry wearing many costumes. The model was never the bottleneck. The absorbing layer is.

None of this is an argument against AI, and this essay is not a nostalgia piece. It is an argument about which side of the pipe is now scarce. Text got infinite. Encoded understanding did not. Every unread answer is a loan, and the interest is paid in re-work.

Exposure is not encoding#

The debt metaphor only holds if exposure and encoding really do come apart. They do, and the literature on it is old, replicated, and brutal. Andy Matuschak made the argument famous in Why books don't work, the book as a medium quietly assumes that people absorb what they read, and almost nobody does. What follows is the measured version of that claim.

Start with what your eyes are doing right now. In studies of mindless reading, the eyes keep moving across the page in a perfectly normal pattern while attention has decoupled entirely (Smallwood 2011). Gaze position says reading. The mind is elsewhere. Across studies, readers spend 20 to 30 percent of reading time mind-wandering, and the more they wander, the less they comprehend (D'Mello & Mills 2021), a correlation around minus 0.3. A fifth to a third of the time you spend reading is exposure without any encoding at all.

Worse, you cannot feel it happening. When learners grade their own recall, the judgment "I got that right" is correct only 57 percent of the time (Dunlosky & Rawson 2012). In the same research program, students who over-trusted that feeling retained 32 percent of the material, and students trained to check themselves against an objective standard retained 89 percent. Same material, same time invested. The difference was calibration. The feeling of understanding is a lagging, lying indicator.

This is why rereading, the world's default study move, keeps losing experiments. In the classic head-to-head, students who reread a passage predicted they would remember it best. A week later the group that had instead spent that time being tested on the passage recalled 56 percent against the rereaders' 42 (Roediger & Karpicke 2006). The rereaders felt better and performed worse, and that inversion is the single most important fact in this entire literature.

The positive version of the fact is called retrieval practice, and it is about as close to settled as this field gets. Three meta-analyses from three different literatures converge. Across 61 lab studies, testing beat restudying at g = 0.50, and the advantage grew as the retention interval got longer (Rowland 2014). Across 217 studies, g = 0.61. Across 48,478 real students in real classrooms, quizzing lifted achievement by half a standard deviation (Yang et al. 2021). Two moderators matter enough to state as rules. Feedback roughly doubles the effect, and without feedback, testing at low success rates does nothing at all, g = 0.03. Retrieval is the spine. Feedback is the vertebrae.

So the mechanics are not in dispute. Memory is built when you produce the material, under some difficulty, with a correction available, and it decays when you merely re-expose yourself to it, however pleasant the re-exposure feels.

NoteCheckpoint

Look away from the screen and say, in one sentence, what this section claimed. If nothing comes back, that is not a failure, it is the section's point demonstrating itself. The claim was, the feeling of understanding and the fact of it are different signals, and only retrieval tells you which one you have.

The graveyard of reading advice#

Before building on the part of the evidence that holds, clear out the part that does not. Reading and learning advice is unusually thick with confident claims that died in replication, and knowing which ones is half the value of doing the research at all.

The claimWhat the evidence saysSource
Match teaching to your learning styleNo credible evidence, after an adversarial review designed to find somePashler et al. 2008
Brain training makes you smarterGains stay on the trained task. The FTC fined Lumosity over its claimsSimons et al. 2016
Working-memory training transfersFar transfer is d ≈ 0.00 with proper controlsMelby-Lervåg & Hulme 2013
Speed reading worksA speed-accuracy trade. RSVP kills the re-reading that repairs comprehensionRayner et al. 2016
Handwriting beats typing for notesFailed replication, pooled g = 0.04. What predicts recall is content captured, not the mediumMorehead et al. 2019, Urry et al. 2021
Highlighting is studyingYour own highlights aid memory a little, g ≈ .36, and comprehension not at allPonce et al. 2022
Expanding review intervals beat fixed onesFolklore. g = 0.034, not significant. Spacing itself is what mattersLatimier et al. 2021
Micro-breaks restore attentionThe famous brief-breaks result failed direct replicationsAriga & Lleras 2011
"ChatGPT boosts learning, g = 0.867"The viral meta-analysis behind it was retracted in April 2026Wang & Fan 2025
"Body doubling works for 85% of ADHDers"The statistic is fabricated. The practice is plausible but has zero controlled trialsEagle et al. 2024

That last pair deserves a beat. A meta-analysis claiming enormous learning gains from ChatGPT went viral, got cited everywhere, and was withdrawn. A body-doubling statistic circulates in every ADHD community and traces back to nothing. Most reading advice is folklore with a citation. The only defense is reading the paper, which is why every number in this essay links to one.

What survived#

Strip away the graveyard and the surviving techniques share one property. Every one of them moves the learner from receiving material to producing it. That is the through-line, and it has a name in the literature, the generation effect, d = .40 across 86 studies, rising to .64 once a day has passed. The rest is engineering around that one fact.

LawThe evidenceIn practice
Retrieve, do not re-exposeThree metas, g = 0.50 to 0.61, growing with delayClose the tab. Say what it claimed. Then check
Space by the horizonOptimal gap is 10 to 20 percent of how long you need it, about 5 percent at a year (Cepeda 2008). Err longReview tomorrow, then next week, then next month
Feedback after every failureWith feedback g = 0.73, without 0.39, and failed retrieval without correction is worth nothingNever end on a wrong answer
Checkpoint through, not afterInterpolated tests cut mind-wandering from 39 to 19 percent and lifted final recall (Szpunar 2013)A question every ten minutes beats a quiz at the end
Guess before you readPretesting beats post-testing, d = 0.30, when reading follows the guess (Pan & Sana 2021). Curiosity lifts recall from 54 to 71 percent (Gruber 2014)The three questions at the top of this essay
Ask exactly what you want to keepTutoring gains collapse from 0.73 to 0.13 when the test does not match the practice (Kulik & Fletcher 2016). Benefits are item-specificWrite the question you actually need answered later
Grade yourself objectivelySelf-graded "correct" is right 57 percent of the time (Dunlosky & Rawson 2012)Type the answer, then compare, no vibes
Self-pace, whatever the modalityReading and listening are equivalent when you control the pace (Clinton-Lisell 2022). Audiobook equalled e-text at two weeks (Rogowsky 2016). Screens lose mainly under time pressure (Delgado 2018)Audio is fine. Skimming against a clock is not
Schedule with a learned modelLearned schedulers beat the 1987 algorithm for 99.6 percent of ten thousand users (SRS Benchmark). Personalized spacing beat generic by a letter grade (Lindsey 2014)Use FSRS, not SM-2, not intuition
Forgive the streakHabits take 66 days at the median, range 18 to 254, and single missed days cost nothing (Lally 2010)A missed day is noise. Punishing it is how systems die

Two of these deserve expansion because they run against instinct.

First, modality does not matter nearly as much as control over pacing. The reading-versus-listening meta found no overall difference, and the small edge that appears for text shows up mainly where readers can regress and pause and listeners cannot. This matters for anyone who reaches for audio or text-to-speech because their eyes slide off dense pages. The evidence says that is a legitimate accommodation, not a lesser way to read, provided the pause button gets used like a tool.

Second, the checkpoint law is the direct answer to the 20-to-30-percent wandering tax. In the interpolated-testing studies, inserting brief retrieval between segments halved mind-wandering and roughly doubled final recall (Szpunar et al. 2013), where merely re-presenting the material did nothing. Structure beats willpower. You do not fix drift by trying harder, you fix it by making the text ask things back.

Produce the memory. Do not receive it. Every law above is that sentence wearing different clothes.

NoteCheckpoint

Same drill as before. Name two laws from the table without scrolling up. Notice how different this feels from the smooth confidence of reading them, that gap is the fluency illusion from section two, experienced live.

You prefer what fails you#

Here is the finding that explains why, if all of this is so well-established, nobody does it and no mainstream product ships it.

In the spacing literature, learners who experienced both schedules and were then shown their own results still got it backwards. 72 percent judged cramming to have worked better even after spacing had objectively improved 90 percent of them (Bjork, Dunlosky & Kornell 2013). The judgment tracks how learning feels while it happens, and effective learning feels worse. Effortful, slow, full of failed retrievals. The field calls these desirable difficulties, and the point of the name is that learners reliably avoid them.

The AI version of this experiment already exists. A 2025 study had people study with AI note-taking at three automation levels. Fully automated notes produced the lowest test scores and the highest preference ratings (2025 preprint). Participants rated the condition that taught them least as best on quality, enjoyment, and intent to use. The best-scoring condition, where AI proposed atomic notes and the human had to assemble and edit them, was liked least. Small study, 30 people, but it lands exactly on the 72-percent result from a completely different literature. Preference and learning are not just uncorrelated. In this domain they point in opposite directions.

Now look at the product landscape with that lens. Summary apps sell compression, read less, know more, which the exam data says is exactly backwards, and the summary category hit its ceiling years ago while subscription learning kept growing around it. Auto-flashcard tools sell the removal of the one effortful step that made the cards work. Beautiful canvas tools sell organization to the already-motivated, a profitable niche that never touches the person whose problem is attention itself. And the tools that do center real retrieval, Anki and its lineage, present like homework software and are abandoned accordingly. The market keeps building vending machines for the feeling of learning because that is what the buyer's thumb votes for.

The one mainstream product that cracked it is instructive. Duolingo runs streaks, leagues, and a learned spacing model over what is, underneath, plain spaced retrieval, and reports a 41 percent daily-to-monthly active ratio (Q1 FY26 letter), retention that consumer apps do not get. Its own research team also published the honest caveat, the famous engagement gains came from better streak mechanics, and the memory model's job was cutting prediction error, not producing the viral number (Settles & Meeder 2016). The lesson is not gamify everything. A meta-analysis of gamification found the cognitive benefit survives rigorous designs while bare points-and-badges effects wash out. The lesson is that adherence machinery around a real retrieval engine is the only combination with receipts.

The market sells the feeling. The learning lives in the friction. Any honest system, personal or product, has to resolve that tension rather than pick a side of it.

The loop#

Everything above compresses into one operating loop. This is the repayment plan for ingestion debt, and it is deliberately boring, six steps, each one carrying a specific finding on its back.

1. Capture into one inbox. Papers into a reference manager, articles into one read-later lane, everything with its source attached. Capturing is cheap and fine. The rule that makes it safe is that capture reserves the option to read, it is not a commitment to read. The collector's fallacy starts where that line blurs.

2. Triage without mercy. For each item, one verdict. Ignore, skim, read, mine for one specific claim, or schedule deep work. The active reading queue stays at five to nine items. Beyond that a queue stops being a plan and becomes guilt storage, and guilt storage is where reading systems go to die.

3. Read with scaffolds, not willpower. Write the purpose before starting, one line, what decision or project does this feed. Guess the answers first, the pretesting effect pays out even when the guesses are wrong. Then checkpoint every ten minutes or every section, look away, state the claim, note what would change your mind. That single habit attacks the mind-wandering tax at its measured source. Dense text through ears instead of eyes is fine, self-pacing is the ingredient that matters (Clinton-Lisell 2022).

4. Convert claims into questions. End every serious read by writing one to five questions, not summaries, questions with checkable answers. This is where AI belongs in the loop, drafting candidate questions from what you highlighted, detecting contradictions against other sources, asking Socratic follow-ups. The one job it must never take is answering the questions for you, because the assembly is the encoding (2025 preprint). Hard cap the volume. Five cards per article, twenty per book chapter. A review queue that outgrows its reader is just the debt again in a new costume.

5. Review on a learned schedule. The questions go into spaced retrieval with FSRS doing the scheduling. Type or say the answer before revealing, then grade against the shown answer, which repairs the 57-percent self-grading problem. Cap daily review time and cut low-value cards without sentiment. Missed a day, nothing happened, continue.

6. Link what survives into notes you can find. One idea per note, source attached, connected to the project it serves. This layer is what makes the whole loop compound, and it only compounds if every writer, human or agent, follows the same shape.

The loop's power-user version has a name in the literature, successive relearning, retrieve to criterion in one session, then again days later, then again. In real courses with real exams it produced gains of half to a full letter grade, d = 0.54 to 1.10 (Janes et al. 2020), one of the largest effects ever measured in situ. The same paper carries the warning that shapes everything I am about to say next. When the practice was optional, 61 percent of students used it zero times. The technique was never the hard part. Capture less. Retrieve more. The hard part is architecture that makes you show up.

What I am building#

I did not run this research pass out of general curiosity. I went in with a product intuition, an infinite canvas of two-minute linked cards for readers whose focus fractures, and I expected the evidence to bless it. It did something more useful. It killed the flattering parts and left the load-bearing ones.

The canvas, it turns out, is commodity. The retention mechanics are the moat. The evidence-mandated shape is a reading companion where the loop itself is the product. It ingests whatever you actually read, segments it around natural stopping points because attention budgets are measured in seconds now, opens each segment with a guess-first question, checkpoints through the text, drafts the retrieval questions for your confirmation, one tap to accept or edit, because user-assembled beats auto-generated and auto-generated is the trap. Then it schedules everything on a learned model and grades recall objectively. The metric it optimizes, and the only one on its dashboard, is what you still know at seven days, because the quiz-day number lies and time-on-page lies harder. The closest prior art is Matuschak's mnemonic medium, essays with the retrieval woven directly into the reading, and the bet here is carrying that pattern from motivated quantum-computing students to the reader whose problem is attention itself.

The audience it serves first is the one every incumbent ignores, the focus-challenged reader. That population is not a niche. 6.76 percent of adults worldwide have symptomatic ADHD (Song et al. 2021), around 366 million people, closer to nine percent in the young-adult cohort, and prescription data shows demand surging, fastest among adult women (CDC 2023). The digital tools marketed at them have, per the umbrella reviews, inconclusive evidence at best. What the learning literature offers instead is scaffolding with receipts, checkpoints, pretesting, forgiving schedules, and that is buildable.

I hold the claims of this section more loosely than the science above, and the open questions are stated in the research, honestly, as A/B tests rather than beliefs. Whether the ideal segment is two minutes or ten. Whether a linear next-step chain helps or fights the curiosity that actually sustains attention. Whether effortful retrieval can be made to feel light enough for the audience that needs it most, which is the make-or-break design problem, because the preference trap does not care how good the evidence is.

Now settle the three guesses from the top. The habit the evidence ranks first is retrieval practice, testing yourself, and if you are like most readers you reread instead. Listening does not retain worse than reading, provided you control the pace. Highlighting barely helps memory and does not help comprehension, though reading someone's provided highlights does, which is what the bold lines in this essay were.

If your three answers were already right, you are ahead of the illusion this essay is about. If any surprised you, that gap between what you believed and what the evidence shows is ingestion debt in miniature, and you now know the repayment mechanics.

Generation is only getting cheaper. The models will keep improving, the answers will keep arriving faster than anyone reads them, and the spread between what passes through you and what stays will keep widening on its own. Retention is the ledger that matters now. Run the loop.

The shelf#

The 44 sources behind this essay, grouped by what they explain, one line each. Every claim above traces to one of these, and they are worth more than my compression of them.

A. Memory and retrieval, 9 sources
  • Rowland 2014, meta-analysis. Testing beats restudy, g = 0.50, growing with delay, worthless without feedback at low success.
  • Adesope, Trevisan & Sundararajan 2017, meta-analysis. g = 0.61 overall, multiple-choice is legitimate, classroom matches lab.
  • Yang et al. 2021, meta-analysis. Classroom quizzing g = 0.499 across 48,478 students.
  • Cepeda et al. 2008, N = 1,354. Optimal review gap is 10 to 20 percent of the retention horizon. Err long.
  • Latimier, Peyre & Ramus 2021, meta-analysis. Spaced retrieval beats massed, g = 1.01. Expanding intervals are folklore, g = 0.034.
  • Brunmair & Richter 2019, meta-analysis. Interleaving helps confusable categories, g = 0.42, not prose.
  • Janes et al. 2020, RCT. Successive relearning lifted real exams 13 percent, d = 0.54 to 1.10. Optional practice went 61 percent unused.
  • Pan & Sana 2021, five RCTs. Guess-first beats quiz-after, d = 0.30, when reading follows the guess.
  • Bertsch et al. 2007, meta-analysis. Generation effect d = .40, rising to .64 past a day.
B. Attention and mind-wandering, 6 sources
C. Modality, screens, audio, notes, 8 sources
D. Metacognition and the myths, 6 sources
E. AI and algorithmic learning, 7 sources
  • Bastani et al. 2025, PNAS, RCT. Unguardrailed GPT help, practice +48 percent, exam −17. Guardrails neutralize, do not add.
  • Wang & Fan 2025, meta-analysis, retracted April 2026. The viral ChatGPT g = 0.867 number is withdrawn.
  • Kulik & Fletcher 2016, meta-analysis. Intelligent tutoring g ≈ 0.50 to 0.66, collapsing to 0.13 on unaligned tests.
  • Kosmyna et al. 2025, preprint. 83 percent of LLM-assisted writers could not quote their own essay. Directional.
  • Lindsey et al. 2014, classroom RCT. Personalized spacing beat generic by 8 to 10 percent, d ≈ 0.9 to 1.4. Massed won the quiz and lost the exam.
  • Settles & Meeder 2016, Duolingo. Trainable forgetting curves cut prediction error 45 percent. The famous engagement gain came from streak mechanics.
  • SRS Benchmark, open dataset. Learned schedulers beat SM-2 for 99.6 percent of about ten thousand users, roughly 350 million reviews.
F. ADHD and the market, 8 sources
  • Lally et al. 2010, longitudinal. Habits form in about 66 days, range 18 to 254. Single missed days are harmless.
  • Gabarron et al. 2025, umbrella review. Digital ADHD interventions, inconclusive due to low-quality evidence.
  • Kollins et al. 2020, RCT. EndeavorRx moved its trained attention metric, none of the clinical secondaries.
  • Song et al. 2021, meta-analysis. Adult symptomatic ADHD, 6.76 percent, about 366 million adults.
  • CDC MMWR 2023. Stimulant fills surged 2020 to 2021, fastest among young adult women.
  • Sailer & Homner 2020, meta-analysis. Gamification's cognitive effect survives rigorous designs, bare points-and-badges effects do not.
  • Eagle et al. 2024, survey. Body doubling, widely practiced, zero controlled trials, the viral statistic fabricated.
  • Duolingo Q1 FY2026 shareholder letter. The habit loop compounds, 41 percent DAU over MAU.

The research itself, five documents and all 44 source summaries with limitations and replication notes, lives in my vault, and this essay is its step six, the synthesis that gets published. If you read all of this and remember one sentence a week from now, let it be the one you produced yourself at a checkpoint. That is the whole theory, demonstrated.