Mathematical Musings

PISA 2025 · A neutral read

The agents’ briefs

The instructions each AI agent received, as reviewed before the agents ran.

Contents

Contents

Each agent starts with no prior context: it sees only its prompt and the files handed to it. Nothing below identifies who commissioned the work. Agents run in this order, since later ones depend on earlier outputs: A first, then B and C and D and E in parallel (B and E get A’s data files), then F on the assembled draft.

Shared instructions prepended to every prompt:

You are producing one section of a neutral analytical report on the PISA 2025 results (released by the OECD on 8 September 2026). “Neutral” means: state what the evidence establishes, what it leaves contested, and what is speculation, and label each claim accordingly. Do not give every claim equal weight; give each the weight the evidence supports. Do not advocate for any policy. Prefer primary sources (OECD reports, technical documentation, national assessment agencies, peer-reviewed research) over commentary; when you use commentary, say so. Cite every factual claim with a source URL or a table number. Where a number has a confidence interval or a statistical-significance flag in the source, report it. Where two sources disagree, report both. If you cannot verify something, say that rather than guessing. Write in plain prose, not bullet lists, except for data tables. Do not use headings that presume a conclusion.


Agent A: Build the dataset and establish the facts#

Materials provided: the PISA 2025 Volume I PDF (377 pages), its extracted text, and the list of StatLink URLs found in it (each downloads an Excel workbook of underlying tables).

Tasks. Download every StatLink workbook and the Annex B1 tables. From them, build a single tidy dataset (CSV) with one row per country/economy × PISA cycle (2000 through 2025) × domain (reading, mathematics, science), containing the mean score, its standard error, the 10th, 25th, 75th, and 90th percentiles where available, the share of students below Level 2 and at Level 5 or above, and the OECD’s flag for whether the trend comparison is valid for that country and cycle. Add the socio-economic gap (difference between top and bottom ESCS quarters) and the gender gap where the tables provide them. Add each country’s PISA coverage index (Coverage Index 3) and school and student response rates for 2018, 2022, and 2025 from Annex A2. Note explicitly which countries the OECD itself flags for sampling or comparability problems in any of these cycles, and what the flag says.

Then write a factual summary of what the data show, with numbers and confidence intervals, covering: OECD-average trends by domain over the full series; the distribution of country-level changes 2022–2025 and 2018–2025 (how many significantly up, down, unchanged); changes at the top and bottom of the within-country distribution, as distinct from changes in the mean; changes in the low-performer share; changes in the socio-economic gap and what is driving them; gender and immigrant-background trends. Reproduce the OECD’s headline numbers and check them against the tables; report any discrepancy. Do not interpret causes.

Separately, read Annex A1, Annex A2, Annex A3, and the Reader’s Guide, and write a short technical note on what changed in the 2025 cycle that could affect comparability: test design (including any adaptive testing), mode, the science framework revision, the linking procedure, the treatment of the 2022 cycle’s pandemic-era samples, and anything else the OECD itself flags. Report what the OECD says about each and whether it presents evidence.

Deliverables: the CSV, a data dictionary, a Python script that regenerates the CSV from the downloaded files, the factual summary, and the technical note.

Agent B: Country trajectories#

Materials provided: Agent A’s CSV, data dictionary, and technical note.

Task. Using the dataset, classify every country with valid trend data by trajectory, using explicit numerical rules that you state (for example: change 2012–2018, change 2018–2022, change 2022–2025, each classified as significant rise, no significant change, or significant fall, with the threshold used). Report the groups. Then, for each group, describe what the member countries have in common and what they do not, drawing on publicly documented system features: GDP per capita, spending per student, share of immigrant students, centralized versus decentralized governance, length of pandemic school closures (use the UNESCO school-closure tracker or OECD data), age of tracking, and reported device-use policies. Where a feature is shared across a group, check whether it is also present in the other groups; the purpose is to identify features that discriminate between trajectories, not features that are merely common. State plainly where no discriminating feature is found. Include the 2000–2025 series for a set of frequently discussed countries (at minimum Finland, Sweden, Poland, Estonia, Germany, the Netherlands, the United Kingdom, Ireland, Singapore, Japan, Korea, Chinese Taipei, Australia, Canada, the United States, Türkiye, and the participating Chinese jurisdictions), and note for each whether the OECD flags any comparability problem. Pay particular attention to systems that rose while most fell, and to whether their rises are real (significant, not driven by sampling changes) before describing them. Produce data tables as an appendix.

Agent C: Claims in circulation#

Task. Collect the explanations for the PISA 2025 results that have been publicly offered since 8 September 2026 and in the run-up to the release, by ministers, journalists, think tanks, academics, and prominent commentators, across as many countries as you can cover (at minimum the United States, United Kingdom, Germany, France, the Netherlands, the Nordic countries, Australia, Canada, Japan, Korea, and Singapore). For each explanation, record who said it, where, the exact claim, and what evidence if any they offered. Do not evaluate the claims; this is an inventory. Group duplicates. Also collect the OECD’s own explanations from the Volume I text and from Andreas Schleicher’s public statements, kept distinct from third-party claims. Include claims that the results are not meaningful or that PISA is flawed, with the same detail. Output a structured inventory with sources.

Agent D: What PISA measures and how far it can be trusted#

Task. Write a review of what PISA does and does not measure, and of the scholarly criticism of it, sufficient for a reader to judge how much weight the results can bear. Cover: the assessment framework and how it differs from curriculum-based assessments such as TIMSS; the psychometric model and the published critiques of it (for example Kreiner and Christensen on the Rasch model; the debates about item-country interaction and rank instability); sampling and coverage issues, including exclusion rates and the coverage index, and the specific 2022 concerns about several countries’ samples; student test-taking effort and the evidence that it varies by country and over time and affects scores; mode effects from the shift to computer-based and adaptive testing; the linking of scales across cycles and how robust trend estimates are; and the broader critiques from Zhao, Sjøberg, Meyer and Benavot, Hopfenbeck, and others, along with the OECD’s responses. For each issue, state what is established, what is contested, and how large the effect could plausibly be relative to the reported changes. Also summarize the evidence on how well PISA scores predict later individual and national outcomes. Do not conclude that PISA is either trustworthy or untrustworthy in general; give the reader the components of that judgment.

Agent E: Candidate causal explanations#

Materials provided: Agent A’s CSV and factual summary, Agent B’s trajectory groups, Agent C’s inventory of claims, and Agent D’s review.

Task. For each candidate explanation of the observed trends, including every distinct one in Agent C’s inventory and any others in the research literature, proceed in three steps and keep them visibly separate in the text.

Step one: state the explanation precisely and write down, before consulting the data, what pattern it would predict. Which countries should be affected more and less; which cycles; which domains; which parts of the score distribution; which socio-economic and demographic groups; what should have happened to the same students’ scores on other assessments.

Step two: check each prediction against Agent A’s dataset, Agent B’s groupings, and other assessment series (national assessments, TIMSS, PIRLS), and report which predictions hold, which fail, and which cannot be tested with available data.

Step three: summarize the peer-reviewed and technical-report evidence on the explanation independent of PISA, with sources, and give a considered assessment of how much of the observed change the explanation could account for, with the reasoning shown.

The explanations to cover at minimum: pandemic school closures and learning loss; smartphone and social-media use and digital distraction; declining reading for pleasure and changes in reading habits; changes in the composition of the student population (immigration, language background, demographic shifts in school-age population); changes in curriculum, standards, and pedagogy (including reform movements in specific countries, treated country by country rather than by slogan); teacher supply, qualifications, and turnover; changes in school accountability and testing regimes; declining student effort or motivation on low-stakes tests; changes in PISA’s own instrument, mode, and sampling (drawing on Agent D); educational spending and austerity; school absenteeism; mental health and wellbeing trends; and secular trends in cognitive test performance (the reversal of the Flynn effect). Where the PISA student questionnaire allows a within-dataset check (device use, sense of belonging, reading enjoyment, absenteeism), perform it and say clearly what a cross-sectional correlation in that data can and cannot show. End with a comparison across explanations: which are consistent with the overall pattern, which are contradicted by it, which are consistent with some countries but not others, and which cannot be distinguished with existing evidence. Do not select a favorite.

Agent G: United States annex#

Materials provided: Agent A’s CSV, factual summary, and technical note; Agent E’s causal review.

Task. Write a United States annex. First, the US PISA record from 2000 to 2025 in all three domains, with confidence intervals, position relative to the OECD average, and the OECD’s sampling notes for each US cycle (US school response rates have historically been below the PISA standard; report the figures and the adjudication outcome each time). Second, place the PISA series alongside the NAEP Long-Term Trend for age 13 and age 17, the main NAEP for grades 8 and 12, and TIMSS grade 8, over the same period; report where the series agree and where they diverge, and discuss what the divergences imply, noting differences in what each assessment measures and whom it samples. Third, for each explanation in Agent E’s review that has been applied to the United States specifically, report what the US-specific evidence shows, including state-level evidence where state NAEP series allow it. Treat explanations attributing US trends to specific standards or curriculum reforms exactly as any other: state the prediction (timing, grades, subjects, states adopting versus not adopting), then check it, with sources. Fourth, note what the US student-questionnaire data show on device use, absenteeism, and reading enjoyment relative to other countries. Produce data tables as an appendix.

Agent F: Adversarial review of the assembled draft#

Materials provided: the full assembled draft report and annex, and Agent A’s CSV.

Task. You are reviewing this report for bias and omission, not for style. Read it as the most skeptical reader from each of several standpoints would: someone who believes the declines are largely a measurement artifact; someone who believes they are real and caused by phones; someone who believes they are real and caused by pandemic closures; someone who believes they are caused by curriculum or pedagogical reform; someone who believes they reflect demographic change; and someone who believes PISA rankings are used to push a particular policy agenda. For each standpoint, identify where the report’s framing, ordering, word choice, or selection of evidence would strike that reader as tilted, and where the report omitted evidence that reader would cite. Separately, check ten numerical claims chosen at random from the report against the CSV or the cited source and report any errors. Check whether any explanation received markedly more or less space than the evidence for it warrants. Output a list of specific, located findings, each with a suggested fix, and an overall judgment of where the report leans if it leans anywhere.