Extract, organize, point to sources
Turn narrative fields into a structured brief, preserve uncertainty, attach source references and flag missing evidence.
A proposed experiment on presentation effects and AI-assisted evidence preparation.
We will use public reconstruction records from Ukraine to test whether the way a project is described changes expert ratings. We will then compare AI-assisted, human-verified briefs with manually prepared briefs, measuring factual accuracy and preparation time.
Data collection pilot completed. Reviewer experiment not yet run.
Each reviewer sees one version of each project.
Each description → the same factual template
Decision-makers have limited attention. Clearer writing may change how the same evidence is judged. We test that possibility.
Reconstruction projects compete for scarce attention and resources. If presentation changes a review score while the underlying facts stay fixed, part of that score reflects how the evidence is communicated.
That could disadvantage communities with less capacity to prepare documents. This is a motivation to test, not a finding we already have.
The practical decision is simple: should a review team keep narrative descriptions, use a fixed factual template, or invest in AI-assisted preparation?
The scale makes careful review consequential: Ukraine’s recovery and reconstruction needs were estimated at almost $588 billion over ten years, as of December 2025. World Bank / joint RDNA5, February 2026 ↗
school roof leaks. 300 pupils. repair the roof. estimated cost UAH 4 million. design documents are available. planned duration 6 months. independent damage assessment not reported.
Same facts. Same missing information. Would you give it the same rating?
Illustration only: these are authored examples, not a DREAM record, an AI run or an experimental result.
Public records supply the material. A controlled reviewer experiment supplies the evidence about wording.
Start with a defined civilian project category. Record each fact, its source and its missing information. Two people check the source evidence; disagreements are resolved.
Output / source-linked dossierCreate plain and polished descriptions from that dossier. Preserve claims, numbers, uncertainty and omissions. Use one language. Check every version before review.
Only intended change / presentationAssign versions to reviewers using a fixed rubric. A reviewer sees only one version of each project and is blind to the treatment. Human reviews are the main test.
Outcome / experimental priority scoreCompare manual and AI extraction into the same template from the same source packet. Measure factual errors, omissions and preparation time—including human corrections—and test the resulting briefs.
Decision / does AI add value?The average change in a project’s review score caused by presentation, with an uncertainty interval.
A review score is an experimental outcome. It does not measure funding awarded, construction quality or social impact.
A rating of a project-version by a reviewer. Preregister one anchored 0–100 score for priority to proceed to further technical review. Average contrasts within projects; account for dependence within projects and reviewers. Report the absolute effect and confidence interval.
Define the eligible project category, sampling date and exclusions first. Use a feasibility pilot of approximately 30 manually checked projects to estimate variance and workload. Choose the main sample through a power analysis, including reviewer and project clustering. Thirty is not a powered final sample.
Recruit people with relevant project-review experience; recruitment is still pending. Use balanced random assignment. Keep project identity and AI labels hidden. Run model reviewers separately with fixed prompts and recorded versions. AI scores cannot stand in for human behavior.
Process each narrative independently with the same frozen AI workflow; the AI does not see the paired version or checked answer sheet. Audit against that sheet afterward. Measure the plain–polished score difference again after normalization and compare it with the raw difference. Identical manually supplied briefs erase wording differences by construction; that alone would not establish an AI benefit.
Every cost needs a currency, unit and source. Keep plans separate from completed work. Mark an unsupported claim as unverified and missing data as “not reported.” Log all failed outputs and corrections; do not silently drop AI errors from the workflow benchmark.
Use Ukrainian for the primary local-review study if the reviewer pool permits it. Treat English as a separate replication, with translation checked by a bilingual reviewer. Generalize only to the chosen projects and reviewers; a national or actual-funding claim needs additional evidence.
These figures describe stored records collected on 12 August 2026. They are not live national statistics.
| Source | What we can use | What it cannot establish |
|---|---|---|
| DREAM ↗Primary project material | Project IDs, descriptions, objectives, document metadata and separate financial records. A documented open API is available; collection should migrate to that route. | A public project record is not necessarily a competitive grant application. A document link is not a verified technical assessment. |
| Diia Digital Community ↗Optional contextual extension | Community scores and official KATOTTG codes for linkage. Useful for exploring differences in record completeness. | The stored response lacks the score’s measurement period. Digitalization is not a validated measure of staffing, writing skill or administrative capacity. |
| New reviewer experimentEvidence we still need to create | Random assignments, ratings, reviewer time, source-fidelity checks and the cost of preparing each brief. | These data do not exist yet. Reviewers must be recruited and the protocol fixed before the main experiment. |
Save dated source copies; deduplicate IDs; inspect currency, units and conflicting values; join by official codes; manually verify each experimental dossier. Unresolved records stay flagged.
Budget, approval, commitment, contract and payment are different observations. They are not a guaranteed sequence and must not be summed as one “funding success” measure. No separate cash-received field was found in the pilot. No reported value means unknown, not zero.
The stored sample establishes data availability. Amounts, dates and supporting evidence still need case-level checks. The study’s primary outcome will come from independently collected expert ratings under randomized presentation, with financial records providing source context.
Download the detailed data audit ↓A fixed template is the baseline. AI earns its place only if it helps fill that template faithfully and reduces total preparation effort.
Turn narrative fields into a structured brief, preserve uncertainty, attach source references and flag missing evidence.
Confirm facts against the public source; inspect contradictions and attachments when needed. AI does not certify engineering quality, truth or eligibility.
Count unsupported additions and omissions, correction time and review-score changes. Compare with manual extraction into a fixed template. A template may be sufficient.
AI writing, proposal review and presentation effects have already been studied. Our proposed contribution is a reconstruction-specific test with checked facts, human reviewers and a non-AI baseline.
Horizon Europe applications already provide evidence on AI writing and evaluation. AI-assisted proposals are not a new research topic.
Structured changes to grant proposals have already been used to test LLM reviewers. An AI-only wording experiment would be too close to this work.
Content-preserving rewrites of manuscripts have also been tested with AI reviewers. This proposal extends that question to reconstruction records, human reviewers and evidence preparation.
This targeted review supports a proposed contribution, not a claim to be first. A fuller literature review and the professor’s assessment should precede preregistration. Read the research review ↓
Estimate how much wording changes experimental scores. This gives review teams a reason to test a common evidence format.
Deliver a source-linked brief template, error audit and estimate of preparation effort that a partner could test in practice.
A precise null result can support a simpler workflow. Wide uncertainty means more evidence is needed—not that an effect is absent.
Start with a defined civilian project category, a reviewed set of about 30 dossiers and a small human-review pilot. Use it to settle the rubric, recruitment feasibility and sample size for a preregistered experiment.