Skip to content
r2RESEARCH
UKRAINE
Full proposal ↗
← One-minute overview
Research design / from data to decisions

How to produce a verifiable economic result.

Fix the budget and the objective, compare transparent rules, account for uncertainty, and calculate the maximum justified price of additional information.

01

Need

Mathematics proficiency in PISA

02

Allocation

Equal capacity, different rules

03

Budget

Information costs reduce the number of places

04

Decision

Compare coverage and justified spending

One clear objective:
reach students in need.

The primary criterion is the share of students below basic mathematics proficiency who are offered a place. Additional criteria describe the composition of recipients and their share of the initial skill deficit.

The target population is limited to students of PISA age represented by the survey in 17 covered Ukrainian regions. Weights do not recover unknown outcomes for excluded territories or children outside the survey’s coverage.

War provides the context and the resource-allocation problem. This design does not separate its causal contribution from pre-war differences, the pandemic or selection into the survey.

Change the rule.
Keep the number of offers fixed.

The main scenario offers places to 20% of the represented population. Capacities of 10% and 30% are checked separately. These are research scenarios, not existing quotas.

A

Random allocation

Every student has the same probability of receiving an offer.

B

Rural priority

Priority for students in a rural school stratum, identified by STRATUM. This is not a verified residential address.

C

Family resources

Priority for the bottom weighted quartile of ESCS among nonmissing values. All ties at the threshold are included.

D

Parental education

Priority for HISCED codes 1–5: highest parental education below short-cycle tertiary education.

How equal capacity is maintained

If the priority group is larger than capacity, each member has the same selection probability. If it is smaller, all members are offered a place and the remainder is allocated randomly among everyone else. We calculate expected coverage rather than draw a lottery of named students.

Students with missing ESCS or HISCED remain in the sample. Missing values do not confer priority under the baseline rule; a separate scenario checks a different treatment.

Skills: 10 plausible values

The basic mathematics proficiency threshold is 420.07. For each of ten plausible values, we separately calculate below-threshold status and the depth of the deficit, max(420.07 − Y, 0).

These values represent uncertainty about the distribution of skills. Averaging them for each child does not produce an exact individual diagnosis.

Sampling: 81 sets of weights

The final weight and 80 replicate weights account for the survey design. The ESCS quartile, priority-group size and selection probabilities are recalculated in every replicate.

We use BRR-Fay with a parameter of 0.5, then add between-imputation variance. Differences between rules are calculated as paired contrasts before pooling.

Uncertainty in the estimateSampling variance + uncertainty about skills

Preliminary intervals: estimate ± 1.96 SE. These do not include future lottery randomness, tutoring effects or unknown properties of a diagnostic test.

Technical specification for the professor

For each plausible value m: Uₘ = 0.05 Σᵣ(θₘᵣ − θₘ)² over 80 replicates. The total is T = mean(Uₘ) + 1.1 × Var(θₘ), where between-imputation variance uses the divisor M − 1; SE = √T.

The original four rules were defined before their pilot results were examined. Missing-data checks, the hybrid rule and capacities of 10%/30% were added after the first calculation and are exploratory. There was no preregistration; the next specification should be fixed before expanding the analysis.

Full methods and exact definitions ↓

Every screening expense
has an alternative use.

The same money could fund more tutoring places. The model determines how much additional spending is offset by a better composition of recipients.

After paying for screeningq(s) = q₀ − s / c

N is population size; c is the cost of a place; B is the budget after common baseline administrative expenses; q₀ = B / (N × c); and s is the additional information cost per person screened.

Everyone is screened, not just recipients

The screening bill is N × s: recipients have not yet been identified when screening is paid for. The baseline model assumes the same cost of reserving a place for every student.

The objective is to preserve at least the baseline number of offers to students below basic proficiency despite the reduction in capacity.

InformationPrice limit relative to the cost of a placeConditions
Perfect knowledge of statuss_max / c = q₀ × (1 − p₀)The below-threshold share exceeds initial capacity. This is only an upper bound.
A test with assumed accuracys_max / c = q₀ × (1 − p₀ / pD)pD > p₀ and there are enough positive results. Otherwise, a separate rule for filling the remaining places is needed.

Here p₀ is the share below the threshold among recipients under the baseline rule. For a test, pD = aL / [aL + (1 − b)(1 − L)], where L is the below-threshold share, a is sensitivity and b is specificity.

Targeting and learning gains are different objectives

If tutoring has the same effect on every student, more costly screening reduces total learning gains by reducing the number of places. Prioritizing students with insufficient skills has an independent rationale as an access objective. Maximizing returns requires data or explicit assumptions about treatment effects.

Completed

Data and reproducibility

The original file has been obtained. Fields, missing values, weights and strata have been checked. Main estimates and standard errors were independently reproduced with separate code.

Completed

Pilot and scenario model

Four rules; checks of capacity, missing information and a hybrid. Spending limits calculated for perfect information and a grid of assumed test accuracy.

Next stage

Robustness of conclusions

Alternative thresholds, leaving out regions one at a time, unequal costs per place and participation scenarios. Fix the primary specification.

Next stage

The main paper

Update the literature review, identify where policy choices are robust and define limits on applying them to real programs. Interviews are not required for the core paper.