Random allocation
Every student has the same probability of receiving an offer.
Fix the budget and the objective, compare transparent rules, account for uncertainty, and calculate the maximum justified price of additional information.
Mathematics proficiency in PISA
Equal capacity, different rules
Information costs reduce the number of places
Compare coverage and justified spending
The primary criterion is the share of students below basic mathematics proficiency who are offered a place. Additional criteria describe the composition of recipients and their share of the initial skill deficit.
The target population is limited to students of PISA age represented by the survey in 17 covered Ukrainian regions. Weights do not recover unknown outcomes for excluded territories or children outside the survey’s coverage.
War provides the context and the resource-allocation problem. This design does not separate its causal contribution from pre-war differences, the pandemic or selection into the survey.
The main scenario offers places to 20% of the represented population. Capacities of 10% and 30% are checked separately. These are research scenarios, not existing quotas.
Every student has the same probability of receiving an offer.
Priority for students in a rural school stratum, identified by STRATUM. This is not a verified residential address.
Priority for the bottom weighted quartile of ESCS among nonmissing values. All ties at the threshold are included.
Priority for HISCED codes 1–5: highest parental education below short-cycle tertiary education.
If the priority group is larger than capacity, each member has the same selection probability. If it is smaller, all members are offered a place and the remainder is allocated randomly among everyone else. We calculate expected coverage rather than draw a lottery of named students.
Students with missing ESCS or HISCED remain in the sample. Missing values do not confer priority under the baseline rule; a separate scenario checks a different treatment.
The basic mathematics proficiency threshold is 420.07. For each of ten plausible values, we separately calculate below-threshold status and the depth of the deficit, max(420.07 − Y, 0).
These values represent uncertainty about the distribution of skills. Averaging them for each child does not produce an exact individual diagnosis.
The final weight and 80 replicate weights account for the survey design. The ESCS quartile, priority-group size and selection probabilities are recalculated in every replicate.
We use BRR-Fay with a parameter of 0.5, then add between-imputation variance. Differences between rules are calculated as paired contrasts before pooling.
Preliminary intervals: estimate ± 1.96 SE. These do not include future lottery randomness, tutoring effects or unknown properties of a diagnostic test.
For each plausible value m: Uₘ = 0.05 Σᵣ(θₘᵣ − θₘ)² over 80 replicates. The total is T = mean(Uₘ) + 1.1 × Var(θₘ), where between-imputation variance uses the divisor M − 1; SE = √T.
The original four rules were defined before their pilot results were examined. Missing-data checks, the hybrid rule and capacities of 10%/30% were added after the first calculation and are exploratory. There was no preregistration; the next specification should be fixed before expanding the analysis.
The same money could fund more tutoring places. The model determines how much additional spending is offset by a better composition of recipients.
N is population size; c is the cost of a place; B is the budget after common baseline administrative expenses; q₀ = B / (N × c); and s is the additional information cost per person screened.
The screening bill is N × s: recipients have not yet been identified when screening is paid for. The baseline model assumes the same cost of reserving a place for every student.
The objective is to preserve at least the baseline number of offers to students below basic proficiency despite the reduction in capacity.
| Information | Price limit relative to the cost of a place | Conditions |
|---|---|---|
| Perfect knowledge of status | s_max / c = q₀ × (1 − p₀) | The below-threshold share exceeds initial capacity. This is only an upper bound. |
| A test with assumed accuracy | s_max / c = q₀ × (1 − p₀ / pD) | pD > p₀ and there are enough positive results. Otherwise, a separate rule for filling the remaining places is needed. |
Here p₀ is the share below the threshold among recipients under the baseline rule. For a test, pD = aL / [aL + (1 − b)(1 − L)], where L is the below-threshold share, a is sensitivity and b is specificity.
If tutoring has the same effect on every student, more costly screening reduces total learning gains by reducing the number of places. Prioritizing students with insufficient skills has an independent rationale as an access objective. Maximizing returns requires data or explicit assumptions about treatment effects.
The original file has been obtained. Fields, missing values, weights and strata have been checked. Main estimates and standard errors were independently reproduced with separate code.
Four rules; checks of capacity, missing information and a hybrid. Spending limits calculated for perfect information and a grid of assumed test accuracy.
Alternative thresholds, leaving out regions one at a time, unequal costs per place and participation scenarios. Fix the primary specification.
Update the literature review, identify where policy choices are robust and define limits on applying them to real programs. Interviews are not required for the core paper.