# Reproducible pilot: allocating 20% of catch-up places, PISA 2025

Run from the repository root:

```sh
outputs/research-2/taxi/pilot-venv/bin/python outputs/research-2/policy-research/pilot/run_pilot.py
```

Alternatively supply `--data /absolute/path/CY09_MS_STU_PUF.sav`. Dependencies are Python, NumPy and pyreadstat; exact versions and source SAV SHA-256 are recorded in `pilot-results.json`. The script reads the original student SAV, selects `CNTRYID == "804"` as a string, exports aggregate estimates only, and never writes student records. The public archive is [OECD CY09_MS_STU_PUF.zip](https://webfs.oecd.org/pisa2022/2025/CY09_MS_STU_PUF.zip). PISA 2025 context and coverage are documented in [the Ukraine country note](https://www.oecd.org/en/publications/pisa-2025-results-volume-i-country-notes_2d4ff9ea-en/ukraine_0eb7d4d3-en.html).

## Population and interpretation

This is a pilot for the enrolled target-age population represented by the PISA 2025 Ukrainian sample in 17 covered regions. It is not a national estimate for every Ukrainian child, every region, refugees abroad, or children outside the assessment's coverage. It is a repeated-survey microdata analysis, not a school panel, experiment, causal effect of war, or evaluation of tutoring benefits. The PISA student weight sum is a survey population total, not a census count.

The policy capacity is a hypothetical 20% of the represented student population. Each rule assigns fractional lottery probabilities, interpreted as expected inclusion in the service. No finite lottery is drawn. Therefore estimates describe expected baseline need among recipients and remaining students, not a realized service cohort. No estimated tutoring effect, cost, earnings return or welfare benefit enters the calculation.

## Fixed rules

1. **Random all:** every student has selection probability 0.20.
2. **Rural first:** prioritize students whose actual `STRATUM` value label contains `Rural/`. All observed labels must classify uniquely as rural or urban; the script stops otherwise.
3. **Bottom ESCS quartile first:** among nonmissing ESCS, compute the weighted lower inverse-empirical-CDF 25th percentile. All observations at or below the cutoff, including exact ties, form the priority group. Missing ESCS remains in the nonpriority group. The priority group is not asserted to contain exactly 25% of the full population or precisely 25% of valid ESCS when there are ties.
4. **Parent nontertiary first:** prioritize valid `HISCED` codes 1–5; codes 6–9 and missing values remain in the nonpriority group. Thus this means highest parental education below short-cycle tertiary under the file's own labels. The code records observed codes and all labels for review.

If priority group weighted share is f ≥ q, its members receive probability q/f and others zero. If f < q, priority members receive probability one and remaining members receive (q−f)/(1−f). Consequently every rule fills exactly q=0.20 of expected population capacity. The missing-ESCS and missing-HISCED groups are retained explicitly. Under a rule whose priority group exceeds capacity, missing cases receive zero probability; this is a substantive consequence of the specified rule, not silent case deletion. It is not a recommendation for handling missing eligibility records in a real program.

For each of 80 replicate weights the code recomputes the ESCS cutoff, ties at that cutoff, priority-group weighted mass and probabilities. Other priority-group definitions stay fixed while their weighted masses and probabilities change. Rules depend only on circumstances, so they are common to all ten plausible values within a weight set.

## Estimands

The primary subject is mathematics, with `PV1MATH` through `PV10MATH`, separately analyzed. The Level 2 cutoff z=420.07 is given in [OECD PISA reporting scales, Annex 17.A](https://www.oecd.org/content/dam/oecd/en/publications/reports/2024/03/pisa-2022-technical-report_599753f0/01820d6d-en.pdf). Below Level 2 means PV < z; deficit is D=max(z−PV,0), measured in fixed PISA score points. Plausible values are not averaged for an individual before thresholding.

With weight w and selection probability π, each statistic is computed independently for every PV and final/replicate weight combination:

- **Selected population share:** Σwπ / Σw, exactly 20% by construction.
- **Low-performance population share:** Σw·1(PV<z) / Σw; this common baseline is repeated in each policy row for convenience.
- **Low-performer recall:** Σwπ·1(PV<z) / Σw·1(PV<z), the fraction of below-Level-2 students reached in expectation.
- **Selected low-performer precision:** Σwπ·1(PV<z) / Σwπ, the below-Level-2 fraction among expected recipients.
- **Deficit share covered:** ΣwπD / ΣwD, the share of the initial aggregate score deficit assigned a place. “Covered” means a place is offered, not that the deficit is eliminated.

All JSON/CSV estimates are proportions on a 0–1 scale. Multiply by 100 for percent; multiply differences and their SEs by 100 for percentage points. Baseline low-performance prevalence and selected capacity are shared across rules and have exactly zero paired differences. Random-all recall and deficit coverage are exactly q, independent of PV and survey sample; their zero SE follows this identity, not precise knowledge about an actual lottery.

## Sampling and plausible-value uncertainty

For PV m, let θ_m use the final weight, and θ_mr use replicate r. With R=80 and Fay factor f=0.5:

```
U_m = [1 / (80 × (1−0.5)^2)] × Σ_r (θ_mr − θ_m)^2
    = 0.05 × Σ_r (θ_mr − θ_m)^2
θ_bar = mean_m θ_m
U_bar = mean_m U_m
B = Σ_m (θ_m − θ_bar)^2 / (10−1)
T = U_bar + (1 + 1/10) × B
SE = sqrt(T)
```

The implementation uses all 10×81 estimates for each metric. For policy-minus-random comparisons, the difference is formed within each PV and weight set **before** applying these formulas. This preserves covariance and is not the square root of a sum of two marginal variances. See [OECD preparation and analysis guidance](https://www.oecd.org/en/about/programmes/pisa/how-to-prepare-and-analyse-the-pisa-database.html) and [2025 technical notes](https://www.oecd.org/en/publications/pisa-2025-results-volume-i_73451bc5-en/full-report/technical-notes-on-analyses-in-this-volume_46a7001c.html). The complete 2025 technical report is forthcoming; this pilot implements the established PISA BRR-Fay and multiple-PV approach rather than claiming that forthcoming report was inspected.

Intervals are normal θ_bar ±1.96 SE, unadjusted for multiple comparisons and not clipped to logical bounds. These are pilot uncertainty intervals. They exclude the extra randomness of an implemented lottery, program effects, program costs, imperfect screening, and unobserved population exclusions. Nonlinear quantile rules can have nonsmooth sampling behavior; recalculating them within every replicate is necessary but does not establish all finite-sample properties.

## Outputs and built-in checks

- `pilot-results.json`: source hash, runtime versions, missingness, cutoff/tie convention, final priority-group probabilities, aggregate estimates, and paired differences.
- `policy-estimates.csv`: four rules × five outcomes, with sampling/MI variance components and normal intervals.
- `paired-differences.csv`: four paired comparisons with random-all × five outcomes.

Assertions verify country coding, 7105 unique students, all ten complete math PVs, all 81 complete weight columns, nonnegative weights, unique rural/urban classification, valid HISCED observed codes, probabilities in [0,1], exact weighted capacity, finite/bounded estimates, and analytical identities for random assignment.

This pilot cannot by itself determine the monetary value of diagnostic testing. A later diagnostic scenario would need an explicit screening accuracy model and cost, evaluated against these implementable-information benchmarks. Perfect latent-skill prioritization is only an unattainable information bound; assigning places using a student's PISA plausible value would not represent a feasible diagnostic test.

## Exploratory sensitivity extension, 12 September 2026

These checks were specified **after** inspecting the four-rule q=.20 pilot. They are exploratory, not preregistered. Run:

```sh
outputs/research-2/taxi/pilot-venv/bin/python outputs/research-2/policy-research/pilot/run_sensitivity.py
```

The script rereads original SAV columns, calculates the same 10×81 design, reproduces all original q=.20 point estimates and variance components, and adds capacities 10%/30%, two missingness policies, and one hybrid policy. It outputs `sensitivity-results.json`, `sensitivity-estimates.csv`, and `sensitivity-paired-contrasts.csv`. This is a computational consistency check, not an independent validation of the estimator. Independent validation is handled separately by the main research task.

For missingness sensitivity at q=.20, missing ESCS or HISCED joins the corresponding priority group instead of the nonpriority group; no student is deleted in either case. For the hybrid at q=.20, rural students receive probability `.25*q + .75*q/f_rural`, and urban students `.25*q`. It is a mixture of expected inclusion probabilities: 25% of capacity is allocated using an all-student lottery rule and 75% using the rural-priority rule. It is not two independent finite lotteries with potentially duplicate recipients. Rural mass exceeds 20% in every final/replicate weight set (range 22.64%–24.47%); every probability is valid. The 25% mixing parameter was not optimized and is not claimed to be socially optimal.

The tables below report percentages. Brackets in comparisons are paired 95% normal intervals for percentage-point differences, formed before BRR/PV pooling.

| Priority rule, q=20% | Below-Level-2 recall | Below-Level-2 precision | Initial deficit share assigned a place |
|---|---:|---:|---:|
| Random all | 20.00 | 45.96 | 20.00 |
| Rural first | 25.31 | 58.17 | 28.83 |
| Bottom valid ESCS quartile first | 25.43 | 58.44 | 28.48 |
| Parent nontertiary first | 23.39 | 53.74 | 24.47 |
| Bottom ESCS or ESCS missing first | 27.29 | 62.71 | 32.02 |
| Parent nontertiary or HISCED missing first | 25.20 | 57.92 | 27.69 |
| Hybrid 25% open / 75% rural | 23.99 | 55.12 | 26.62 |

Rural-minus-ESCS paired differences are −0.116 percentage points for recall [−2.221; 1.989], −0.265 for precision [−5.111; 4.581], and +0.352 for deficit coverage [−3.345; 4.049]. These estimates do not establish a difference between the two specified rules, and do not demonstrate equivalence. Comparing overlaps between the separate rule intervals would not answer this question.

Including missing ESCS in priority increases recall by 1.857 percentage points [1.267; 2.447] and initial deficit coverage by 3.542 [2.493; 4.590]. The corresponding HISCED changes are +1.817 [1.228; 2.407] and +3.219 [2.355; 4.082]. Missing covariates are associated with greater observed need in this sample. This supports further study of data-collection and eligibility exclusion; it does **not** establish the cause of missingness, justify treating missing answers as poverty, or recommend rewarding omitted documentation. A real policy would require a fair way to obtain or review missing eligibility information and an assessment of response incentives.

The hybrid assigns urban students a 5% inclusion probability and rural students 68.54%, against 0% and 84.72% under pure rural priority. Its recall falls 1.329 percentage points [−1.769; −0.888] and initial deficit coverage falls 2.208 [−2.923; −1.493] compared with pure rural priority. Against the all-student lottery its deficit coverage is higher by 6.625 [4.479; 8.770]. This quantifies one open-access/need-incidence tradeoff; it does not establish that geographic access is the correct definition of fairness.

| Capacity | Rural recall / deficit share | ESCS recall / deficit share | Parent education recall / deficit share |
|---|---:|---:|---:|
| 10% | 12.66 / 14.42 | 12.72 / 14.24 | 11.69 / 12.24 |
| 20% | 25.31 / 28.83 | 25.43 / 28.48 | 23.39 / 24.47 |
| 30% | 35.75 / 39.55 | 35.75 / 38.98 | 35.08 / 36.71 |

At 30% capacity rural and ESCS priority groups are fully covered, with remaining places assigned randomly to the other students. Parental nontertiary priority still exceeds capacity. The CSV/JSON contains every marginal SE and paired contrast with random allocation at the new capacities. All these quantities concern places offered against initial need, not skills restored.
