# Cognitive Debt and the Measurement of Reasoning Autonomy in the Age of Artificial Intelligence

**A construct definition and validation architecture for the Cognitive Autonomy Index**

Neuro ASI / Phroneme · Draft for review · 2026

---

## Abstract

Large language models have made it possible to produce competent reasoning artifacts, essays, analyses, arguments, without performing the reasoning oneself. Early neurophysiological evidence indicates that this substitution is not cost-free: when people offload a cognitive task to an AI assistant, the neural systems that would otherwise perform the work engage less, and the reduced engagement can persist after the assistance is removed. We take this finding seriously without overreading it. We argue that the socially and economically important variable is not whether AI is "good" or "bad" for cognition, but whether an individual retains the capacity to reason without assistance, and that this capacity is currently unmeasured. We define that capacity as **reasoning autonomy** and propose the **Cognitive Autonomy Index (CAI)** as its operational measure. We specify the construct, distinguish it from general intelligence, describe a measurement architecture built on named item response models, rotating equated forms, and an explicit treatment of the reliability of change, with an optional neural convergent-validity anchor, and lay out a falsifiable validation design. We are explicit throughout about what the instrument cannot claim. The measurement rests on the parts of the brain sciences that replicate, and it is designed around, not on top of, the parts that do not.

## 1. Introduction

The dominant public question about artificial intelligence and the mind is framed badly. "Is AI making us smarter or dumber" invites a verdict that the evidence cannot support and that will date with the next news cycle. The better question is narrower and measurable: as capable assistance becomes ambient, does a given person retain the ability to reason when the assistance is removed, and can that ability be tracked over time.

There is now a first data point suggesting the question is not rhetorical. In a 2025 study from the MIT Media Lab, fifty-four participants wrote essays under three conditions, unaided, with a search engine, and with a large language model, while their brain activity was recorded with electroencephalography (Kosmyna et al., 2025). The group that used the language model showed the weakest measured connectivity across networks associated with planning, integration, and recall, and a majority could not quote a sentence from the essay they had just produced. When participants who had relied on the model were later asked to write unaided, the reduced-engagement pattern did not immediately reverse. The authors named the residue **cognitive debt**. Survey and behavioral work points in a compatible direction, reporting an inverse association between frequent AI tool use and critical-thinking performance that is statistically mediated by cognitive offloading (Gerlich, 2025).

We do not treat one EEG study as settled neuroscience, and Section 6 states plainly why no such study should be. We treat it as a signal that a real variable exists, that the variable has begun to move, and that no one is measuring it at the level of the individual over time. That gap is the subject of this paper.

## 2. Standing on the floor, not the ceiling

Any measurement of cognition inherits the epistemic condition of the brain sciences, which is uneven. Claims sitting close to molecules and cells are solid. The Hodgkin and Huxley model of the action potential has held for seventy years (Hodgkin and Huxley, 1952). Claims sitting close to a thought, a feeling, or a diagnosis are far weaker. Roughly thirty-participant task-based neuroimaging studies replicate only modestly (Turner et al., 2018); when seventy independent teams analyzed a single functional imaging dataset against nine preregistered hypotheses, no two chose the same analytic pipeline and their conclusions diverged materially (Botvinik-Nezer et al., 2020); and robust brain-behavior correlations appear to require sample sizes in the thousands rather than the dozens (Marek et al., 2022). The thirty-year public story that depression reflects a serotonin deficit was found, on systematic review, to rest on no consistent evidence (Moncrieff et al., 2022).

The strategic consequence for a measurement company is specific. An instrument that claims to read reasoning off a brain scan is building on the contested ceiling. An instrument that measures reasoning behaviorally, under standardized conditions, and uses neural data only as convergent evidence rather than as the measurement itself, is building on the floor. The CAI is designed for the second position.

## 3. The construct: reasoning autonomy

We define **reasoning autonomy** as the capacity to carry a novel problem from premises to a defensible conclusion without external cognitive assistance, under standardized conditions, as demonstrated by performance rather than self-report.

Three commitments follow from the definition.

**It is a capacity, not a trait, and not general intelligence.** The CAI does not attempt to rank innate ability. It measures whether a specific, trainable capacity is currently intact and whether it is trending up or down. This distinction is not cosmetic. A trait framing invites the entire disparate-impact liability of the intelligence-testing tradition; a capacity-and-trajectory framing keeps the instrument developmental, which is both more honest and more defensible.

**It is behavioral first.** The measurement is performance on reasoning tasks, not the activation of any region and not a questionnaire. Self-report scales carry an assumed validity that the wider field has not earned, so the CAI does not lean on them.

**It is longitudinal by design.** A single value is nearly meaningless and is easy to misuse. The unit of analysis is the trajectory, the change in reasoning autonomy across repeated standardized measurements, which is where both the scientific signal and the individual value live.

**Why offloading does not build the capacity.** The construct is defined by what a person can do unaided, and the reason unaided performance responds to practice rather than to exposure is one of the better-replicated findings in the learning literature. Retrieval is the operative act. When people study a prose passage and are then asked to recall it, they retain more of it a week later than people who reread the same passage for the same time, even though rereading looks more effective when the test comes minutes after study (Roediger and Karpicke, 2006). The effort of generating an answer is not friction around learning; it is the event that produces the retention. This bears on AI assistance in a specific way. Producing a reasoning artifact with a capable assistant is a procedural skill, and it can be learned well, but what is learned is the operation of the tool. It does not require the retrieval and inference the same task would have demanded unaided, and it does not require building the intermediate representation of the problem that a person needs in order to carry the reasoning to an unfamiliar case. The structural version of the point is unremarkable and predates the current tools: an agent that never acts and never receives feedback on the consequences of its own action has no error signal to learn from (Sutton and Barto, 2018). That framework is computational and is not a claim about human neural mechanism, but the structure transfers cleanly enough to be worth stating. The practical consequence is that fluency at prompting can rise while reasoning autonomy stays flat or falls, and nothing except a measurement of the unaided case can tell those two apart.

## 4. Measurement architecture

The instrument is a bank of reasoning tasks calibrated with item response theory (IRT), delivered as rotating, statistically equated parallel forms under adaptive selection, with an optional neural convergent-validity anchor, and reported trajectory-first.

**The measurement model is named, not gestured at.** Dichotomous reasoning items are calibrated with the two-parameter logistic model, which estimates a difficulty and a discrimination for each item (Embretson and Reise, 2000); partial-credit and constructed-response items use Samejima's graded response model (Samejima, 1969). Dimensionality is tested before calibration rather than assumed, because reasoning autonomy is a hypothesis about structure and not a license to sum arbitrary items; where the data are multidimensional, a bifactor or multidimensional model is used. Item fit and local independence are evaluated for every item, and misfitting items are revised or retired before administration. The three-parameter guessing term is admitted only where a genuine pseudo-guessing floor exists, because it estimates unstably and often launders weak items. The working demonstration on this site runs exactly this scoring family, offline, so the machinery can be inspected rather than merely described.

**Equated parallel forms and adaptivity.** Because any consequential measure invites people to train for the test rather than for the underlying capacity, a fixed form would decouple the score from the construct it is meant to track. This is Goodhart's law, and it is the default fate of consequential metrics, not an edge case. The defense is to remove the fixed target: a continuously refreshed item bank, equated through a common-item nonequivalent-groups anchor design and placed on a single scale by concurrent calibration (Kolen and Brennan, 2014), delivered adaptively by maximum information so that no two administrations are identical, with item-exposure controls, content balancing, and item-parameter-drift monitoring that quarantines any item whose behavior shifts.

**The trajectory is the product, so the reliability of change is the central problem.** A difference between two noisy measurements is noisier than either, and this is where most longitudinal claims quietly fail (Cronbach and Furby, 1970). The instrument therefore never reports a bare difference. Each occasion carries an ability-conditional standard error; a change is called real only when it exceeds a minimum detectable change and a reliable change index clears its threshold (Jacobson and Truax, 1991); and the trajectory itself is estimated as a latent slope in a growth or latent-change-score model, with individual slopes shrunk toward the population so that a single lucky session is not read as a trend. The levers on change reliability, more information per occasion, optimal spacing between administrations, and pooling across occasions, are treated as the core engineering problem rather than a footnote.

**Longitudinal measurement invariance is a precondition, not a courtesy.** A trajectory across rotating forms is interpretable only if the instrument measures the same construct the same way over time (Meredith, 1993). Invariance is established in sequence, configural then metric then scalar then strict, and only the levels that hold are used to license comparison; where full invariance fails, partial-invariance models anchor on the items that hold and the limitation is reported. Without at least metric and scalar invariance, ordinary drift is indistinguishable from cognitive decline, which is the most common invalid longitudinal claim in this area.

**Neural anchor as convergent validity, not measurement.** Optionally, and with the consent architecture in Section 5, EEG or fNIRS may be recorded during a subset of administrations. This neural signal is never the score. It serves two narrow purposes: convergent validity, testing whether behavioral CAI movement corresponds to the engagement changes reported in the offloading literature (Kosmyna et al., 2025), and a difficult-to-clone credibility anchor. Because these connectivity metrics carry their own modest test-retest reliability, convergent correlations are disattenuated for the anchor's unreliability rather than read at face value. Treating neural data as convergent evidence rather than as the measurement is the precise inverse of the reverse-inference error described in the companion paper.

**Aggregate-first reporting.** For institutional use the load-bearing deliverable is the de-identified cohort pattern, not the individual ranking. This is a measurement choice and an ethics choice at once, and Section 5 explains why.

## 5. Validation and the ethical architecture

**A falsifiable Phase 0.** The first study is a construct-validity study with preregistered, disconfirmable predictions: that CAI scores show acceptable test-retest reliability across equated forms; that they correlate with established reasoning measures while remaining distinct from a pure speed or vocabulary factor; that they move in the expected direction under experimental manipulations of cognitive offloading, net of the practice gains that repeated testing introduces; and that differential item functioning across demographic groups, tested with Mantel-Haenszel and an item-response likelihood-ratio method, is reported and used to remove biased items. A construct that cannot specify what result would falsify it is not a measurement, it is a brand.

**Misuse designed out, not contracted away.** An instrument that ranks reasoning is, in law, a cognitive test, and the moment a score touches an admissions, employment, or accommodation decision it enters the disparate-impact and accessibility law that has bankrupted more sophisticated organizations than a startup can survive. A license clause forbidding such use does not bind third parties and does not survive a subpoena. The controls are therefore architectural. The individual owns the score and can delete it. No individual score reaches an institution without explicit, revocable consent. Institutional reporting is aggregate by default, which starves the individual-harm predicate that a disparate-impact claim requires. Neural data is treated as the highest-sensitivity tier, minimized, firewalled, and governed by an explicit retention and destruction schedule, because biometric-privacy statutes punish collection mechanics, not only misuse.

**Honesty as compliance and as position.** Publishing the methodology, including its confidence intervals and its failure modes, is simultaneously a scientific norm, a regulatory shield against the overclaim that cost a prior brain-training company a federal settlement (FTC, 2016), and a competitive position that a marketing-led rival structurally cannot copy.

## 6. Limitations

We hold the following limits in view rather than in a footnote.

**Reverse inference bounds the neural anchor.** Activation in a region does not establish that a specific process occurred, because most regions participate in many processes (Poldrack, 2006). The neural anchor can therefore corroborate behavioral movement but can never become the measurement.

**Change is harder to measure than level.** The reliability of a difference score is bounded below the reliability of either measurement, so a trajectory instrument lives or dies on the change machinery in Section 4 (Cronbach and Furby, 1970). We report a minimum detectable change and model the trajectory as a latent slope precisely because a raw difference would overstate certainty.

**A trajectory assumes invariance that must be earned.** Comparing scores across rotating forms and across time is licensed only after configural, metric, and scalar invariance are established (Meredith, 1993). Until then the instrument reports within-form results and treats cross-time comparison as provisional.

**Emotion and cognition do not localize cleanly.** Meta-analysis shows discrete mental categories do not map onto dedicated regions; the amygdala, the popular "fear center," is not fear-specific (Lindquist et al., 2012). Any neural interpretation the instrument surfaces must be probabilistic and bounded.

**Reproducibility discipline is mandatory.** Given the analytic flexibility documented in the imaging literature (Botvinik-Nezer et al., 2020) and the sample sizes real effects require (Marek et al., 2022), CAI validation must preregister pipelines and report effect sizes with intervals, not point estimates dressed as certainty.

**Goodhart is permanent, not solved.** The equated-form and adaptive machinery manages gaming; it does not end it. Vigilance is a running cost.

**Individual variability is large.** Population-level findings are weak guides to a single person, so the instrument reports trajectories and confidence, not verdicts.

**The narrative is perishable.** If the culture stops worrying about cognitive offloading, demand framed on that worry evaporates. The construct, verified reasoning capacity, retains value under fear, scarcity, or compliance, so the instrument is anchored to the durable asset rather than to the current anxiety.

## 7. Conclusion

Cognitive debt is the first measurable sign that ambient assistance changes the reasoning brain. The response is not a manifesto about AI. It is an instrument: a behaviorally grounded, longitudinally reported, ethically architected measure of the capacity to reason without help, built on the parts of the science that hold and honest about the parts that do not. That honesty is not a concession. In a field littered with overclaim, it is the most defensible position available, and it is the one an Ivy committee and a serious investor will both recognize.

---

## References

1. Botvinik-Nezer, R., Holzmeister, F., Camerer, C. F., et al. (2020). Variability in the analysis of a single neuroimaging dataset by many teams. *Nature, 582*(7810), 84-88.
2. Cronbach, L. J., and Furby, L. (1970). How we should measure "change": Or should we? *Psychological Bulletin, 74*(1), 68-80.
3. Embretson, S. E., and Reise, S. P. (2000). *Item Response Theory for Psychologists.* Lawrence Erlbaum Associates.
4. Federal Trade Commission (2016). Lumosity to Pay $2 Million to Settle FTC Deceptive Advertising Charges. FTC Press Release.
5. Gerlich, M. (2025). AI Tools in Society: Impacts on Cognitive Offloading and the Future of Critical Thinking. *Societies, 15*(1), 6.
6. Hodgkin, A. L., and Huxley, A. F. (1952). A quantitative description of membrane current and its application to conduction and excitation in nerve. *The Journal of Physiology, 117*(4), 500-544.
7. Jacobson, N. S., and Truax, P. (1991). Clinical significance: A statistical approach to defining meaningful change in psychotherapy research. *Journal of Consulting and Clinical Psychology, 59*(1), 12-19.
8. Kolen, M. J., and Brennan, R. L. (2014). *Test Equating, Scaling, and Linking: Methods and Practices* (3rd ed.). Springer.
9. Kosmyna, N., Hauptmann, E., Yuan, Y. T., et al. (2025). Your Brain on ChatGPT: Accumulation of Cognitive Debt when Using an AI Assistant for Essay Writing Task. MIT Media Lab (preprint, arXiv:2506.08872).
10. Lindquist, K. A., Wager, T. D., Kober, H., Bliss-Moreau, E., and Barrett, L. F. (2012). The brain basis of emotion: A meta-analytic review. *Behavioral and Brain Sciences, 35*(3), 121-143.
11. Marek, S., Tervo-Clemmens, B., Calabro, F. J., et al. (2022). Reproducible brain-wide association studies require thousands of individuals. *Nature, 603*(7902), 654-660.
12. Meredith, W. (1993). Measurement invariance, factor analysis and factorial invariance. *Psychometrika, 58*(4), 525-543.
13. Moncrieff, J., Cooper, R. E., Stockmann, T., et al. (2022). The serotonin theory of depression: a systematic umbrella review of the evidence. *Molecular Psychiatry, 28*, 3243-3256.
14. Poldrack, R. A. (2006). Can cognitive processes be inferred from neuroimaging data? *Trends in Cognitive Sciences, 10*(2), 59-63.
15. Roediger, H. L., and Karpicke, J. D. (2006). Test-enhanced learning: Taking memory tests improves long-term retention. *Psychological Science, 17*(3), 249-255.
16. Samejima, F. (1969). Estimation of latent ability using a response pattern of graded scores. *Psychometrika Monograph Supplement, 34*(4, Pt. 2).
17. Sutton, R. S., and Barto, A. G. (2018). *Reinforcement Learning: An Introduction* (2nd ed.). MIT Press.
18. Turner, B. O., Paul, E. J., Miller, M. B., and Barbey, A. K. (2018). Small sample sizes reduce the replicability of task-based fMRI studies. *Communications Biology, 1*, 62.
