Research proposal · Ukraine

Evaluating Ukraine’s Reconstruction Projects

A proposed experiment on presentation effects and AI-assisted evidence preparation.

We will use public reconstruction records from Ukraine to test whether the way a project is described changes expert ratings. We will then compare AI-assisted, human-verified briefs with manually prepared briefs, measuring factual accuracy and preparation time.

Data collection pilot completed. Reviewer experiment not yet run.

The experiment at a glance

Public DREAM project recordResearcher-checked facts
A
Plain description
Same needSame costSame evidence
B
Polished description
Same needSame costSame evidence
01 / Independent human reviewsDoes presentation change the score?

Each reviewer sees one version of each project.

02 / Compare brief preparation

Each description → the same factual template

Manual preparationResearcher extracts facts
vs
AI + human verificationAI extracts; researcher checks
Review-score differencesErrors & omissionsTotal time
Proposed design · No experimental results yet
The economic problem

Decision-makers have limited attention. Clearer writing may change how the same evidence is judged. We test that possibility.

A strong description
is not the same as
a strong project.

Reconstruction projects compete for scarce attention and resources. If presentation changes a review score while the underlying facts stay fixed, part of that score reflects how the evidence is communicated.

That could disadvantage communities with less capacity to prepare documents. This is a motivation to test, not a finding we already have.

The practical decision is simple: should a review team keep narrative descriptions, use a fixed factual template, or invest in AI-assisted preparation?

The scale makes careful review consequential: Ukraine’s recovery and reconstruction needs were estimated at almost $588 billion over ten years, as of December 2025. World Bank / joint RDNA5, February 2026 ↗

One project · three ways to read itInvented teaching example
School roof repair

school roof leaks. 300 pupils. repair the roof. estimated cost UAH 4 million. design documents are available. planned duration 6 months. independent damage assessment not reported.

=

Same facts. Same missing information. Would you give it the same rating?

Illustration only: these are authored examples, not a DREAM record, an AI run or an experimental result.

Hold the facts still.
Change the presentation.

Public records supply the material. A controlled reviewer experiment supplies the evidence about wording.

  1. 1

    Build a checked fact sheet

    Start with a defined civilian project category. Record each fact, its source and its missing information. Two people check the source evidence; disagreements are resolved.

    Output / source-linked dossier
  2. 2

    Write two faithful versions

    Create plain and polished descriptions from that dossier. Preserve claims, numbers, uncertainty and omissions. Use one language. Check every version before review.

    Only intended change / presentation
  3. 3

    Randomize independent reviews

    Assign versions to reviewers using a fixed rubric. A reviewer sees only one version of each project and is blind to the treatment. Human reviews are the main test.

    Outcome / experimental priority score
  4. 4

    Compare preparation workflows

    Compare manual and AI extraction into the same template from the same source packet. Measure factual errors, omissions and preparation time—including human corrections—and test the resulting briefs.

    Decision / does AI add value?
Δ
What we will actually estimate

The average change in a project’s review score caused by presentation, with an uncertainty interval.

A review score is an experimental outcome. It does not measure funding awarded, construction quality or social impact.

For the professor: identification, sample and measurement

Unit & primary outcome

A rating of a project-version by a reviewer. Preregister one anchored 0–100 score for priority to proceed to further technical review. Average contrasts within projects; account for dependence within projects and reviewers. Report the absolute effect and confidence interval.

Sample & assignment

Define the eligible project category, sampling date and exclusions first. Use a feasibility pilot of approximately 30 manually checked projects to estimate variance and workload. Choose the main sample through a power analysis, including reviewer and project clustering. Thirty is not a powered final sample.

Reviewer population

Recruit people with relevant project-review experience; recruitment is still pending. Use balanced random assignment. Keep project identity and AI labels hidden. Run model reviewers separately with fixed prompts and recorded versions. AI scores cannot stand in for human behavior.

A fair test of the remedy

Process each narrative independently with the same frozen AI workflow; the AI does not see the paired version or checked answer sheet. Audit against that sheet afterward. Measure the plain–polished score difference again after normalization and compare it with the raw difference. Identical manually supplied briefs erase wording differences by construction; that alone would not establish an AI benefit.

Facts & uncertainty

Every cost needs a currency, unit and source. Keep plans separate from completed work. Mark an unsupported claim as unverified and missing data as “not reported.” Log all failed outputs and corrections; do not silently drop AI errors from the workflow benchmark.

Language & external validity

Use Ukrainian for the primary local-review study if the reviewer pool permits it. Treat English as a separate replication, with translation checked by a bilingual reviewer. Generalize only to the chosen projects and reviewers; a national or actual-funding claim needs additional evidence.

Yes, for a bounded pilot.
Not for every claim.

These figures describe stored records collected on 12 August 2026. They are not live national statistics.

100public project records collectedSelected by most recent update
98/100have an exact community-code matchAmbiguous multi-community links need review
85/100have a recorded forecast costPresence does not certify the amount
0human or AI evaluation runsNo measured presentation effect yet
Data sources and their role in the proposed study
SourceWhat we can useWhat it cannot establish
DREAM ↗Primary project materialProject IDs, descriptions, objectives, document metadata and separate financial records. A documented open API is available; collection should migrate to that route.A public project record is not necessarily a competitive grant application. A document link is not a verified technical assessment.
Diia Digital Community ↗Optional contextual extensionCommunity scores and official KATOTTG codes for linkage. Useful for exploring differences in record completeness.The stored response lacks the score’s measurement period. Digitalization is not a validated measure of staffing, writing skill or administrative capacity.
New reviewer experimentEvidence we still need to createRandom assignments, ratings, reviewer time, source-fidelity checks and the cost of preparing each brief.These data do not exist yet. Reviewers must be recruited and the protocol fixed before the main experiment.

Collect → check → link → freeze

Save dated source copies; deduplicate IDs; inspect currency, units and conflicting values; join by official codes; manually verify each experimental dossier. Unresolved records stay flagged.

Inspect the pilot & audit ↗
Why funding records are not the primary outcome

Budget, approval, commitment, contract and payment are different observations. They are not a guaranteed sequence and must not be summed as one “funding success” measure. No separate cash-received field was found in the pilot. No reported value means unknown, not zero.

The stored sample establishes data availability. Amounts, dates and supporting evidence still need case-level checks. The study’s primary outcome will come from independently collected expert ratings under randomized presentation, with financial records providing source context.

Download the detailed data audit ↓

An assistant to
prepare evidence.
A hypothesis to test.

A fixed template is the baseline. AI earns its place only if it helps fill that template faithfully and reduces total preparation effort.

AI

Extract, organize, point to sources

Turn narrative fields into a structured brief, preserve uncertainty, attach source references and flag missing evidence.

Human

Check every claim

Confirm facts against the public source; inspect contradictions and attachments when needed. AI does not certify engineering quality, truth or eligibility.

Study

Measure benefit and failure

Count unsupported additions and omissions, correction time and review-score changes. Compare with manual extraction into a fixed template. A template may be sufficient.

If qualified human reviewers cannot be recruited, the fallback is a smaller study of model behavior and extraction reliability. It would support a narrower claim.

A specific contribution.
An honest boundary.

AI writing, proposal review and presentation effects have already been studied. Our proposed contribution is a reconstruction-specific test with checked facts, human reviewers and a non-AI baseline.

Real reconstruction records+Human review experiment+Accuracy & preparation cost
The closest prior work—and how this proposal differs
2026 · Research preprint

Same substance, different academic rhetoric ↗

Content-preserving rewrites of manuscripts have also been tested with AI reviewers. This proposal extends that question to reconstruction records, human reviewers and evidence preparation.

This targeted review supports a proposed contribution, not a claim to be first. A fuller literature review and the professor’s assessment should precede preregistration. Read the research review ↓

A useful answer.
Even if AI does not help.

If presentation matters

Quantify the sensitivity

Estimate how much wording changes experimental scores. This gives review teams a reason to test a common evidence format.

If AI improves the workflow

Provide a tested protocol

Deliver a source-linked brief template, error audit and estimate of preparation effort that a partner could test in practice.

If effects are small or absent

Avoid an unnecessary tool

A precise null result can support a simpler workflow. Wide uncertainty means more evidence is needed—not that an effect is absent.

The proposal to take to a professor

Approve a feasibility study.
Make the larger claim earn its evidence.

Start with a defined civilian project category, a reviewed set of about 30 dossiers and a small human-review pilot. Use it to settle the rubric, recruitment feasibility and sample size for a preregistered experiment.

Agree on the project categorySecure access to reviewersApprove the protocol & power plan
Open the one-page brief