# Research review: presentation, verification, and reconstruction project review

**Review date:** 10 September 2026. **Decision:** continue with a substantially narrower question. This memo reviews the repository and a targeted set of primary academic and official sources; it is not an exhaustive systematic literature review. No experiment has been run by this review.

## Recommendation in plain language

The useful question is whether reviewers judge the **same reconstruction project differently because of its presentation**, and whether a source-linked standard brief can reduce that difference without losing information. This is a feasible study of information and administrative burden. The present data cannot establish that AI closes Ukraine's administrative capacity gap or changes actual funding allocations.

Study title: **Evaluating Ukraine’s Reconstruction Projects**. Subtitle: **A proposed experiment on presentation effects and AI-assisted evidence preparation.**

Suggested main question: **When the facts of a civilian reconstruction project are held constant, how much does presentation change human reviewers' priority assessments, and can an AI-assisted, human-verified brief reduce that sensitivity at an acceptable verification cost?** AI reviewers are a separate comparison group, not a substitute for evidence about human decisions.

The practical output is a tested review format and an estimate of when AI helps to prepare it. An equally useful outcome would be that a simple template works just as well, or that verification costs make the AI workflow unattractive.

## Why the question matters, and why the current framing is too broad

Unequal local administrative resources are a documented reason to care about project preparation. The OECD's 2022 review reports that 87% of surveyed urban municipalities considered themselves sufficiently staffed to identify investment needs and prepare proposals, compared with 66% of rural municipalities. These are historical self-reports, not an estimate of today's presentation bias. [OECD, *Rebuilding Ukraine by Reinforcing Regional and Municipal Governance*](https://www.oecd.org/en/publications/rebuilding-ukraine-by-reinforcing-regional-and-municipal-governance_63a6b479-en/full-report/component-5.html).

DREAM is a relevant setting because it already connects investment planning, project information, and funding processes. Official reporting states that the 2026–2028 Medium-Term Public Investment Plan was developed using DREAM. The study should explain its additional research purpose rather than imply that Ukraine needs another project registry. [DREAM, 9 July 2025](https://www.dream.gov.ua/news/article-100).

The existing website combines at least three papers: digital capacity and project preparation; capacity and actual funding/progression; and a presentation experiment using AI judges. Each requires different evidence. A digital index is also not interchangeable with staffing, engineering expertise, fiscal resources, or administrative capacity. The strongest available first paper is the controlled presentation study; the linked digital index can supply descriptive context and exploratory subgroup analysis if its timing and coverage are validated.

Calling all public DREAM project records “applications” is misleading. An identifiable application to a particular funding call, a public investment project, a portfolio inclusion decision, a contract, and a payment are different observations. A later real-funding study should select a defined programme, application cohort, eligibility rules, decision date, and evaluation rubric.

## What is already known: a defensible contribution

The broad concepts are established. The claim “nobody has done this before” is not supportable.

| Primary source | What it already studies | Implication for this proposal |
| --- | --- | --- |
| [*The effect of writing style on success in grant applications*, Journal of Informetrics, 2022](https://www.sciencedirect.com/science/article/pii/S1751157722000098) | Associations between application language, panel scores, and grant selection, with applicant covariates. | The idea that presentation relates to grant decisions is not new; observational text associations alone do not isolate a causal style effect. |
| [Wiles, Munyikwa, and Horton, *Algorithmic Writing Assistance on Jobseekers' Resumes Increases Hires*, NBER working paper; published in Management Science, 2025](https://www.nber.org/papers/w30886) | A large field experiment in resume writing assistance; treated participants were hired more often, without evidence of lower employer satisfaction. | Writing assistance changing access to an opportunity is already studied. Improved communication can reveal useful information; every presentation effect is not automatically unfair discrimination. |
| [Zheng et al., *Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena*, 2023](https://arxiv.org/abs/2306.05685) | LLM judging, including position, verbosity, and self-preference concerns. | An AI-only demonstration of sensitivity to formatting or verbosity is a limited contribution. Human agreement is task-specific. |
| [Santoleri, Rentocchini, and Lelli, European Commission JRC, *LLM-assisted proposal writing in competitive R&D funding: Evidence from Horizon Europe*, 18 March 2026](https://publications.jrc.ec.europa.eu/repository/handle/JRC146131) | Firm applications to Horizon Europe in 2021–2024; LLM-writing adoption and evaluation/funding associations. Repeated-submission analysis does not show clear evidence that adopting LLM assistance itself worsens evaluations. | AI-assisted writing, applicant disadvantage, and public funding are already connected in close prior work. Distinguish our experimentally fixed facts and reconstruction setting. |
| [Thorne et al., *Evaluating LLM-Based Grant Proposal Review via Structured Perturbations*, March 2026, preprint](https://arxiv.org/abs/2603.08281) | Six EPSRC proposals, controlled perturbations, different LLM review architectures, and human assessment of reviews. | AI grant review and perturbation tests already exist. The proposed study needs an appropriately sampled domain dataset and actual human treatment outcomes. |
| [Hazra et al., *Accepted with Minor Revisions: Value of AI-Assisted Scientific Writing*, ACM IUI 2026](https://arxiv.org/abs/2511.12529) | An incentivized randomized study of abstract editing, authorship disclosure, and human review. | Human review of AI-assisted scientific writing is also an existing research area. |
| [Li et al., *How Can Rhetoric Reward-Hack AI Reviewers?*, 10 August 2026, preprint](https://arxiv.org/abs/2608.08975) | 120 anonymized manuscripts, 4,200 content-preserving variants, and five AI reviewers; rhetorical choices change assessments. | Even the exact general mechanism “same scientific content, different rhetoric, different AI score” is already studied. |

**Candidate contribution:** a source-audited reconstruction-project benchmark that separates presentation from factual content, compares human and AI reviewers, and measures the factual fidelity and labor cost of a practical standardization workflow against a simple template. This combination may be useful and distinctive; the targeted search does not establish uniqueness. A professor should assess whether the sector, reviewer access, and scale make the contribution sufficient for the intended degree or journal.

## What the repository establishes, and what it does not

The checked-in August 12 pilot contains 100 recently updated DREAM records and reports 98 project-to-index matches across 63 primary communities. These figures describe an existing local snapshot. They are not a current national census or a probability sample. See [the original discovery audit](data-discovery.md).

The pilot demonstrates that public records and code-based linkage can be collected. It does not yet establish that the available documents contain enough comparable need, readiness, cost, and output facts for a review experiment. That requires a document-level audit in the chosen sector. A record hosted on an official portal is evidence of what was reported; it does not independently establish the engineering validity of a budget or the truth of a beneficiary forecast.

The website should say **“checked against the source record”** where appropriate. Reserve **“independently verified”** for facts checked against a second authoritative document or qualified technical review.

Specific repository concerns observed during this review:

- `build_pilot.py` constructs the 30-project AI set by sorting sector, PQS, and cost and taking the first 30 rows. It is a construction sample, not random or stratified selection.
- The current T1/T2 prose and T4 JSON do not expose the same facts: T4 includes `funding_gap` and `other_verified_facts`, which are omitted from the prose. T0 is original text and also does not hold the factual set constant. These comparisons cannot be described as pure presentation contrasts.
- Numeric validation is bypassed for T0/T4, and `unsupported_additions: 0` is hard-coded. An identical hash attached to different texts does not validate semantic equivalence. Human review is still pending.
- The factual-core `funding_gap` calculation uses missing identified funding as zero. Missing financing is not evidence of no financing. Keep the gap unknown unless its inputs and interpretation are established.
- Document counts and procurement counts are observable metadata, not an independently validated measure of technical readiness.
- PQS mixes record presence, document metadata, readability, and outcome detail. It should not be presented as a validated measure of project merit or as pure writing quality. Funding can cause records to become more detailed, so current completeness–funding associations also permit reverse causality.
- The source audit and website disagree on the number of digital-index components; the original audit reports four, while one website row says five. Reconcile against the payload before public display.

These observations are tied to the code reviewed on September 10; subsequent fixes should be recorded separately.

## A feasible experiment

### 1. Narrow and freeze the population

Start with one civilian sector and a reasonably comparable project stage, for example educational-facility reconstruction proposals before procurement. Confirm that the available records support this restriction before fixing the sector. Freeze a dated frame, document inclusion/exclusion rules, identify duplicates and multi-community projects, then draw a reproducible random or stratified sample. Do not select projects for having an interesting funding outcome or for extreme text scores.

Use the source language for the main domestic-review study, with reviewers proficient in that language. If the substantive question is donor-facing English review, explicitly select that population and apply the same independently checked translation procedure to every arm. Existing source-provided English text can otherwise confound translation quality with original preparation capacity.

### 2. Build an auditable factual dossier

Two independent coders extract the same predeclared fields: problem/need, proposed intervention, target outputs and units, intended beneficiaries and definition, dated cost with currency and basis, timeline, readiness evidence, risks, and source links. Resolve disagreements before randomization. Preserve “not reported,” conflicting values, forecasts, and uncertainty explicitly.

Every retained claim needs a source location, date, and status such as reported, corroborated, or unresolved. Validate numbers, units, dates, entities, negation, conditions, and missingness. “Capacity: 500 places” must not silently become “500 people served per year.” A text may contain the same numbers and still change its meaning.

### 3. Create and blind the presentation treatments

Create two versions of each dossier: an ordinary, less structured narrative and a polished, clearly structured narrative. Both contain the identical substantive claims and qualifications. Avoid caricatured spelling errors or emotional claims that alter perceived need. Keep length approximately matched; alternatively predeclare length as a separate manipulated feature. Use independent reviewers to confirm that the manipulation changes perceived polish while preserving facts.

Show each human reviewer **only one version of a given project**, under a neutral random label. Randomize assignment within project and reviewer workload blocks, balance order, and conceal treatment labels, project identity, index score, and actual funding outcome. No reviewers involved in a project should assess it. A “within-project” contrast does not mean that the same person sees both versions and recognizes the manipulation.

### 4. Test the remedy without making it tautological

A manageable factorial design is presentation (ordinary/polished) × review input (raw/AI-assisted verified standard brief). Each brief must be independently produced from its assigned input by a frozen pipeline that cannot see the counterpart, the treatment label, or the gold dossier. A separate checker audits it and records all corrections and time. If the intervention includes human correction, describe the result as the performance of **the AI-plus-verification workflow**, not autonomous AI.

Replacing both treatments with one identical gold fact card guarantees removal of the original presentation treatment by construction. That may illustrate a template, but is not experimental evidence that AI can reliably create it. The substantive test is whether the pipeline preserves information, handles missingness, and produces acceptable review inputs without an excessive correction burden.

Benchmark against a **deterministic template populated from existing structured fields**. Also time a small manual, no-AI brief-preparation baseline from the same source packet. Report differences in input coverage honestly: a schema-only template cannot retrieve claims that exist only in free text. AI has an incremental use case only if it improves useful extraction or preparation time after verification relative to these baselines.

### 5. Outcomes and analysis

Use one preregistered primary outcome: **a human priority-for-further-review score on a fixed 0–100 rubric**. Draft the rubric with relevant reconstruction or public-investment experts and pilot its comprehension. Need, expected social benefit, evidence of readiness, cost rationale, and feasibility may be separate dimensions; missing evidence should trigger uncertainty, not invented values.

Secondary outcomes: factual comprehension, correct identification of missing information, reviewer confidence, time per review, and agreement between reviewers. Hypothetical shortlist decisions can be an extension. Neither a score nor an elicited probability is an observed funding decision, and subjective confidence is not calibrated accuracy.

For each reviewer population, estimate the raw presentation difference and the presentation difference after the workflow. The interaction estimates how the workflow changes presentation sensitivity. Use project and reviewer blocking/fixed effects as justified by assignment, and inference that accounts for shared projects and reviewers. If several projects share a community, incorporate that dependence in the power simulation and inference. Do not control for post-treatment comprehension or output quality in the main total-effect estimate.

Run AI reviewers separately with fixed prompts and recorded model versions; use separate contexts, randomized order, and multiple model families. Repeated calls measure stochastic variation. They do not turn 30 underlying projects into hundreds of independent projects. Report human and AI effects separately before any comparison.

A smaller score gap is only useful if the workflow retains decision-relevant facts. Report additions, omissions, source-link accuracy, corrections, unresolved cases, and labor cost alongside scores. Do not silently drop failed standardizations. Preregister how failures enter the operational result and any per-protocol analysis. Avoid “mitigation rate” as the headline because it is unstable when the original score gap is close to zero; show both score differences and uncertainty in points.

### 6. Pilot, power, and recruitment

The existing 30 records are suitable for learning about dossier preparation after reselection or an explicitly labeled convenience pilot. They do not justify a powered final claim. An illustrative workload for **30 projects × 4 conditions × 3 independent human ratings = 360 assessments** could be split among 30 reviewers doing 12 assessments each. This is planning arithmetic, not an evidence-based sample-size recommendation or a claim that reviewers have been recruited.

Use the pilot to estimate factual-review time, eligibility/attrition, reviewer and project variance, treatment reliability, and recruitment feasibility. Decide the smallest useful score change with the professor and practitioners, then simulate power for the actual crossed, clustered assignment. Choose the main sample size only after that exercise and recruit it independently of promising pilot results. Three ratings per cell and a 0–100 scale are design starting points, not validated requirements.

Without access to suitable human reviewers, the honest fallback is an AI-evaluation benchmark. Its conclusion must be limited to those models under the recorded prompts. A later randomized assistance trial with communities or funding programmes is required to estimate real administrative burden, submission success, or allocation effects.

## What belongs on the first page

1. **The issue:** reviewers see project descriptions; preparation resources differ.
2. **One concrete illustration:** one hypothetical school project, two fact-equivalent descriptions, then the same review question. Label it as an illustration with no measured score.
3. **The test:** source documents → audited facts → randomized presentation → human and AI review → comparison.
4. **What AI does:** draft a sourced standard brief, flag missing facts, and reduce repetitive preparation if it passes verification. It cannot repair absent surveys, invent benefits, certify engineering feasibility, or decide what a community deserves.
5. **Data and status:** dated source records; current 100-record feasibility pilot; 30 constructed dossiers awaiting manual validation; no evaluator results.
6. **Value and contribution:** a tested format, measured sensitivity, factual error and time estimates, and evidence about whether AI adds anything over a template.
7. **Decision for a professor:** help select a sector and rubric, assess the contribution relative to close literature, and establish access to qualified reviewers.

Put exploratory regressions, editable metrics, imports, and detailed funding semantics in an appendix. The first page should enable a professor to explain the question, contrast, data, result, and limit after one reading.

## A short professor pitch

**English:** Reconstruction projects are evaluated through documents, and communities differ in the resources available to prepare them. I propose to test whether human reviewers give different priority scores to the same civilian reconstruction project when its presentation changes. Public DREAM records provide the source material; independent coding will fix the facts before randomized, blinded review. We will then test an AI-assisted, human-verified standard brief, benchmark it against a simple template, and measure factual errors and preparation time. The contribution is a practical test of presentation sensitivity in reconstruction review, with human and AI evaluators studied separately. The current 100-record dataset establishes collection feasibility; it contains no experimental result. The immediate next step is a 30-project design pilot with a defined sector, rubric, and reviewer pool, followed by a preregistered study sized using pilot variance.

**По-русски:** Мы хотим проверить, меняется ли оценка одного и того же проекта восстановления только из-за того, как он описан. Берём реальные публичные записи DREAM, проверяем содержащиеся в них сведения и делаем две версии с одинаковыми фактами. Разные эксперты оценивают их вслепую. Затем проверяем, помогает ли единая карточка, подготовленная с AI и проверенная человеком, уменьшить разницу в оценках. Сравниваем её с обычным шаблоном и считаем ошибки и затраченное время. Результатом будет понятный ответ, нужен ли здесь AI и при каких условиях. Сейчас есть проверка доступности данных; доказательств влияния на оценки или финансирование ещё нет.
