Skip to main content

Module A — Experimental Design and Regression Pipeline

1. Purpose of this memo

This memo consolidates the empirical design behind the investment-project analysis. Its goal is not to report final results. Its goal is to define, in operational and scientific terms, both the experiment that the recovered notebooks were attempting to implement and the stronger experimental framework now used to evaluate candidate designs.

The memo answers:

What empirical experiments can we credibly run with the recovered data, what data objects and assumptions define each experiment, and what validity checks must pass before an estimate is interpreted?

A separate memo should define the project annotation and classification protocol. That second memo should own the rules for labeling projects as jobs-related, direct-jobs, indirect-jobs, non-jobs, macro-policy, locally implemented, and so on. This memo assumes those labels are inputs to the empirical pipeline.

The key boundary is now:

Annotation protocol
project_id -> project-level labels

Canonical empirical data
projects + locations + dates/status + geography + outcomes + covariates

Experiment specification
treatment + timing + geography + exposure + counterfactual + outcome + sample rules

Analysis sample
-> validity / measurement gates
-> estimator family
-> falsification and sensitivity
-> research interpretation

This is a deliberate change from treating the empirical pipeline as simply:

project labels -> area-period exposure -> matching -> regression

The recovered matching workflow remains important, but matching is now treated as one estimator family inside a larger experimental system, not as the definition of the research design itself.


Current scientific architecture — 2026 reframe

The active research architecture separates three layers that should not be collapsed.

A — empirical infrastructure

This layer describes the world and preserves provenance. It should remain stable even if the preferred experiment changes.

Examples include:

  • projects and source identifiers;
  • project-level classification labels;
  • geocoded project locations;
  • project timing and status fields;
  • GADM administrative units;
  • GID × TimePeriod panels;
  • ACLED/UCDP outcomes;
  • Afrobarometer observations and geographic links;
  • DHS and population covariates;
  • spatial precision and source metadata.

A project should not be intrinsically labeled “treated” in this layer. Treatment is experiment-specific.

B — experiment specification

An experiment defines the empirical contrast to be studied. At minimum, it should make the following explicit:

treatment definition
project status / treatment timing
geographic unit or exposure radius
counterfactual group
pre-treatment window
post-treatment window
outcome
sample restrictions
estimator family

Different experiments may reuse the same canonical data while making different defensible choices.

C — validity and calibration gates

Before interpreting estimates, each candidate experiment should be checked for:

  • data integrity;
  • timing resolution;
  • treatment/control support;
  • geographic and temporal overlap;
  • outcome coverage;
  • pretreatment balance / selection;
  • geolocation precision;
  • exposure collisions and multiple projects;
  • placebo or falsification behavior;
  • bandwidth or spatial sensitivity where relevant;
  • synthetic signal recovery / detectable-effect calibration.

The operational interpretation is:

A green gate means permission to investigate further, not causal validation.

A failed gate should identify where the experiment fails — data, measurement, support, sensitivity, identification, or substantive signal — rather than merely produce another coefficient.

Estimator family

The current architecture allows multiple estimator families:

Estimator family
├── raw / descriptive comparison
├── matching-based comparison
├── completed vs planned spatial comparison
├── longitudinal / staggered-treatment panel design
└── future spatial-spillover-aware designs

These should be treated as complementary scientific tools. A favorable coefficient from one specification is not a reason to select that specification after the fact.

Methodological lessons from the Blair / Briggs design lineage

Recent methodological reconstruction of Blair, Marty & Roessler (2022), Briggs (2019), and closely related geocoded-aid studies adds several concrete requirements to the FCV design space:

  1. Future or planned project locations can be candidate counterfactuals. They may absorb part of the non-random geography of project placement better than an assumed “never treated” group, but the comparability assumption must be diagnosed rather than assumed.
  2. Project status must be evaluated at the observation date. Planned, active, completed, and ambiguous states should not be collapsed into a timeless treatment flag.
  3. Spatial bandwidth is a scientific parameter. Radius sensitivity is useful, but changing a radius also changes the identifying sample; it is not automatically a clean dose-response test.
  4. Effective identifying N matters more than total row count. Stronger geographic fixed effects or stricter counterfactual definitions can sharply reduce usable support.
  5. Geolocation precision must be preserved. A short computed distance is not highly informative when the underlying source coordinates are coarse.
  6. Exposure collisions must remain visible. Locations may be exposed to multiple donors, projects, sectors, or temporal states.
  7. Selection diagnostics are substantive evidence about the design. Differences between future-project and never-project locations should be reported rather than hidden.
  8. Falsification belongs upstream of interpretation. Pre-outcome placebos, fake timing, alternate exposure rules, or negative controls help determine whether an apparent signal is trustworthy.
  9. Signal recovery should be calibrated before interpreting a null. Synthetic or semi-synthetic effect injection can show whether the empirical apparatus is capable of recovering a plausible weak effect under the observed noise and clustering structure.

These lessons expand the recovered design; they do not invalidate the original matching work.


2. Research question

The broad research question is:

Do African administrative areas exposed to development investment projects — especially employment-relevant or jobs-related projects — experience different post-treatment trajectories in violence, political legitimacy, civic engagement, service delivery, or related outcomes than comparable areas without such investments or with non-jobs-related investments?

The design is motivated by a causal question. The recovered notebooks most clearly implement a matching-based empirical comparison, but the active project should now be described as a family of candidate experiments rather than as one finalized matching design. Causal interpretation depends on treatment definition, timing, project geolocation quality, counterfactual construction, overlap, outcome coverage, falsification behavior, and assumptions about unobserved confounding.

The core empirical idea is now:

  1. Identify and validate development investment projects funded by World Bank and/or Chinese development finance sources.
  2. Classify those projects into treatment-relevant categories, especially jobs-related versus non-jobs-related projects.
  3. Preserve project timing, status, geolocation, and provenance in canonical data objects.
  4. Define a candidate experiment explicitly: treatment, timing, geography/exposure, counterfactual, outcome, and sample rules.
  5. Build the corresponding analysis sample and run validity / measurement gates before estimation.
  6. Estimate the contrast using one or more pre-declared estimator families, including matching where appropriate.
  7. Apply falsification and sensitivity checks before substantive interpretation.
  8. Compare stability across source families, treatment definitions, administrative levels, time windows, outcome families, and estimator families.

3. Unit hierarchy

The analysis uses several nested units. Confusion between these units was one of the main causes of difficulty in the original workflow.

LayerUnitMain role
Raw projectProjectSource metadata: title, objective, sector, amount, dates, country, funding source.
Project-locationProject × geocoded locationSpatial exposure source; one project can have many locations.
Administrative areaGADM L1 / L2 / L3 areaGeographic unit for exposure assignment, covariates, and outcomes.
Area-periodGID × time windowMain recovered analysis-panel unit.
Experiment specificationNamed design objectDefines treatment, timing, geography/exposure, counterfactual, outcome, sample, and estimator family.
Analysis sampleExperiment-specific observationsRows that satisfy the experiment definition and remain after explicit data/support rules.
Matched pairTreated area-period + pure control area-periodDerived object for one matching estimator family.
Matched trioJobs-related area-period + other-investment area-period + pure control area-periodDerived object for a possible multi-arm matching design.

The annotation protocol should label projects at the project level. The empirical pipeline should convert project-level labels into experiment-specific exposures through geocoded project locations, spatial joins or radius rules, and explicit time rules.


4. Empirical estimand and interpretation

The desired empirical quantity depends on the experiment specification. One candidate quantity is the difference in post-treatment outcomes between exposed and comparable unexposed administrative areas.

A generic matched-pair estimand is:

E[Y_post | treated, matched] - E[Y_post | control, matched]

A generic future/planned-location contrast is:

E[Y | implemented/completed exposure] - E[Y | selected future/planned exposure]

A regression representation for the recovered area-period lane is:

Y_{g,t+k} = alpha + beta * Treatment_{g,t} + gamma * X_{g,t-1} + delta_t + epsilon_{g,t}

where:

  • g is a geographic area, such as a GADM L2 or L3 unit;
  • t is the treatment time window;
  • k indexes the post-treatment outcome lag or outcome period;
  • Treatment_{g,t} is exposure to a project category during the treatment window;
  • X_{g,t-1} are pre-treatment covariates;
  • delta_t are time-window controls or fixed effects;
  • beta is the main coefficient of interest for that specification.

The interpretation should remain conservative unless the chosen experiment resolves its main identification issues. Matching, planned/completed comparisons, and panel/event-time designs each rely on different assumptions and should not be described as interchangeable causal estimators.


5. Data lineage

The recovered notebook chain implies the following historical pipeline:

Raw project sources
-> standardized project/location CSV and GeoJSON files
-> project-location to administrative-area assignment
-> area x time-period investment exposure panel
-> merged covariates and outcomes
-> matched samples
-> diagnostics and regressions

The current operating layer should sit on top of that recovered work as:

recovered / refreshed source data
-> canonical empirical entities
-> experiment specification
-> analysis sample
-> validity gates
-> estimator family
-> falsification / sensitivity
-> interpretable research output

5.1 Notebook-stage map

StageExisting notebooks / filesCurrent interpretation
Raw source ingestion50 - Raw data to GDF, 501, 502, 503, 54Ingests World Bank public/project data, AidData World Bank geocoded data, and AidData China/TUFF-style development finance data.
Spatial intersection51 - Spatial IntersectionAssigns project locations to GADM administrative areas.
Investment by admin area52 - Investment by admin area (GID), 55 - Exploration GeographyAggregates project exposure by administrative area and time.
WB/China coexistence diagnostics53 - Chinese and WB coexisting locationsExplores spatial overlap or proximity between Chinese and World Bank projects.
Project classification support56 - Exploration Sectors, 57 - Explore for Job-Related Investments, 58 - WB Projects info to DocSupports manual and rule-based project classification; does not itself finalize labels.
Panel construction80 - Preprocess DataBuilds area-period panels combining investments, violence, population, and DHS geocovariates.
Outcome variables81 - Outcome VariablesConceptual or partial outcome layer; needs consolidation.
Sample and covariate checks82 - Sample Selection Checks, 83 - Covariate ChecksChecks treatment imbalance and covariate availability.
Matching84 - KNN Matching, older 98 - knn MatchingImplements matched treated/control samples, mainly one-to-one matching.
Matching diagnosis85 - KNN Matching DiagnosisEvaluates sample sizes, balance, and matched-pair quality.
Regression prototypes85 - Regression Analysis, 85 - Regression Analysis 2, 99 - OLS regressionsPrototype regression and estimator-debugging notebooks; not yet a final canonical regression module.

6. Source families

The core investment source families are:

Source familyApproximate roleNotes
World Bank public / project dataProject metadata and text fields for classificationUseful for titles, objectives, sectors, themes, lending instruments, project pages, and approval dates.
AidData World Bank geocoded releaseProject-location exposure layerImportant for geocoded locations; may have older time coverage than newer World Bank public project data.
AidData China / TUFF / Global Chinese Development FinanceChinese development finance project exposureImportant comparator or treatment family; geolocation and source-version issues require validation.
GADM administrative boundariesSpatial unit backboneUsed to assign project locations and outcomes to L1/L2/L3 administrative areas.
ACLED / UCDPViolence outcomesUsed for conflict, deaths, fatalities, violence against civilians, and related violence categories.
DHS geocovariatesPre-treatment covariates and possibly outcomesUsed for baseline socioeconomic and household/environmental controls.
AfrobarometerPolitical legitimacy, participation, civic engagement, service deliveryRequired for the non-violence outcome side of the project; coverage and merging need validation.
Population / GHSL or related population dataDenominators and controlsUsed to scale investment exposure and construct per-capita measures.

7. Treatment definitions

The recovered notebooks use basic source-family treatments. The full research design requires additional treatments derived from the annotation protocol.

7.1 Existing source-family treatments

TreatmentArea-period definition
cnwb_pooledArea-period exposed to either World Bank or Chinese project finance.
wb_onlyArea-period exposed to World Bank project finance and not Chinese project finance.
cn_onlyArea-period exposed to Chinese project finance and not World Bank project finance.
any_investmentArea-period exposed to any relevant investment source.
pure_controlArea-period with no relevant investment exposure.

7.2 Treatments requiring annotation

TreatmentArea-period definition
jobs_anyExposure to at least one project labeled jobs_direct or jobs_indirect.
jobs_directExposure to a project explicitly targeting jobs, employment, public works, skills, vocational training, labor-market insertion, or similar direct employment mechanisms.
jobs_indirectExposure to a project plausibly generating employment indirectly through infrastructure, agriculture, private-sector development, education, market access, or productive capacity.
non_jobs_investmentExposure to an investment project not classified as jobs-related.
macro_policy_onlyExposure to or identification of projects that are mainly macro, policy, institutional, budget-support, or technical-assistance interventions without clear local implementation.
locally_implementedProject has identifiable local or physical implementation activities, not merely national-level policy support.

The annotation memo should define these project-level labels. This memo defines only how those labels enter an experiment specification.

Treatment should also preserve temporal state when the source data permit it. A candidate experiment may distinguish:

planned / future
active / ongoing
completed
ambiguous timing
never observed in the project source

These states should be derived relative to the observation date rather than permanently attached to the project.


8. Candidate comparison designs

Eric's email guidance implies several nested comparisons. These should be preserved as candidate experiments derived from the substantive research question, not treated as obsolete simply because the estimator architecture has broadened.

The “output type” column below records the recovered or originally requested matching representation. The same substantive contrast could later be estimated with another defensible design if its assumptions and data support are stronger.

DesignComparisonRecovered / candidate output typeStatus
Pair AAny investment vs pure controlMatched pairsBasic version exists for source-family treatments.
Pair BWorld Bank investment vs pure controlMatched pairsBasic version exists.
Pair CChinese investment vs pure controlMatched pairsBasic version exists, but sparse coverage is a concern.
Pair DJobs-related investment vs pure controlMatched pairsRequires stable annotation labels.
Pair EDirect-jobs project exposure vs pure controlMatched pairsRequires refined labels.
Pair FIndirect-jobs project exposure vs pure controlMatched pairsRequires refined labels.
Trio AJobs-related investment vs non-jobs investment vs pure controlMatched triosConceptually requested; not yet confirmed as canonical implementation.
Trio BDirect jobs vs indirect jobs vs pure controlMatched trios or multi-arm designRequires refined annotation and sufficient sample size.
WB/CN contrastWorld Bank vs Chinese projects vs pure controlMatched trios or separate matched analysesNeeds careful source comparability and sample-size review.

A second candidate counterfactual family should now be investigated where project timing/status supports it:

Design familyComparisonMain identifying ideaStatus
Planned/completed spatial comparisonCompleted or implemented project exposure vs future/planned project exposureCompare locations selected for projects at different implementation states to absorb part of project-location selectionMethodologically promising; FCV source-field feasibility not yet validated.
Longitudinal area-period designArea trajectories before/after treatment timingUse repeated area-time observations and explicit treatment timingParticularly relevant for ACLED/UCDP; requires modern staggered-treatment design choices.

The jobs-related versus non-jobs-related comparison remains central because it is the comparison most closely tied to the substantive theory that employment-relevant development projects may affect violence, legitimacy, participation, or service delivery differently than other investments.


9. Time structure

The recovered pipeline uses time-windowed panels rather than a single static cross-section.

Core time descriptors:

DescriptorMeaning
TLength of time window, usually 2, 3, or 4 years in the recovered notebooks; 1-year windows were discussed as a possible extension.
y0Starting-year convention, commonly 2000 or 2001 in recovered matching outputs.
Treatment windowPeriod in which investment exposure is measured.
Observation-date project statePlanned, active, completed, or ambiguous status evaluated relative to the observation date when source fields permit it.
Outcome windowPost-treatment period in which violence, survey, or service outcomes are measured.
Pre-treatment covariatesCovariates measured before treatment exposure or at baseline.
Lag ruleRule ensuring outcome information is not contemporaneous with or prior to the treatment definition.
Cross-sectional fallbackA possible simplified design if panel windows become too sparse.

Eric explicitly emphasized the need to compare different time-period cutoffs because sample size and outcome coverage may change substantially across window definitions. The final pipeline should therefore treat time-window length and project timing convention as design parameters, not hardcoded choices.


10. Geography

The recovered area-period analysis uses GADM administrative areas, usually within Africa. Survey-linked or project-radius experiments may additionally use direct spatial exposure rules.

GeographyRoleCurrent assessment
GADM L1Large administrative regionsUseful for diagnostics, but likely too coarse for credible matching.
GADM L2Intermediate administrative regionsLikely a plausible main specification for the recovered area-period lane.
GADM L3Finer local administrative regionsPotentially more spatially precise, but may increase sparsity and missing data.
Radius-based exposureRespondent or location within a specified distance of a projectRelevant to Blair/Briggs-style survey linkage; must preserve source geolocation precision and radius sensitivity.
Africa subsetMain continent-scale analysis universeNeeds clear country inclusion and source coverage validation.
Country-level restrictionsPossible robustness or source-coverage filtersTo be defined after source-version validation.

A practical rule for the next phase:

Treat ADM1 as diagnostic; prioritize ADM2 and ADM3 for substantive area-period comparisons subject to sample-size and outcome-coverage checks. Treat spatial radius and geolocation precision as explicit experiment parameters rather than universal preprocessing constants.


11. Outcomes

The empirical design has several outcome families. Not all are equally ready in the recovered pipeline.

Outcome familyExamplesCurrent status
ACLED / UCDP violenceEvents, deaths, fatalities, violence against civilians, battles, conflict categoriesMost clearly integrated in recovered regression/matching prototypes.
AfrobarometerState legitimacy, political participation, social participation, civic engagement, service deliveryExplicitly requested; needs source and merge validation.
DHSHousehold or local socioeconomic proxies; also covariatesDHS geocovariates appear in panel construction; outcome role needs clarification.
Population / GHSLPopulation denominators and scalingUsed for exposure denominators and controls, not main outcomes.
Pre-treatment covariatesNightlights, infrastructure, DHS geocovariates, pre-treatment violence, terrain/forest/roughness where availableNeeded for matching and regression adjustment.

A key unresolved task is to classify each variable as one of:

main outcome
pre-treatment covariate
denominator / scaling variable
diagnostic variable
excluded / not used

Outcome families do not necessarily need the same geographic or estimator architecture. ACLED/UCDP can exploit a dense area-time panel, while Afrobarometer may be better suited to respondent/community spatial exposure and survey-time project status.


12. Matching estimator family — recovered implementation

The later matching notebook implements a one-to-one treated/control matching procedure. This remains an important recovered estimator family, but it is no longer treated as the complete definition of the empirical design.

The logic is:

  1. Load area-period panel data.
  2. Define a binary treatment variable.
  3. Separate covariates from treatment and outcomes.
  4. Group data by time period.
  5. Drop units with missing covariate information.
  6. Compute distances between treated and candidate control units using observed covariates.
  7. Use linear assignment / Hungarian matching to minimize total covariate distance.
  8. Save matched treated-control pairs.
  9. Diagnose balance and sample size.

The matching grid includes:

region: africa
admin level: l1, l2, l3
T: 2, 3, 4 years
y0: 2000, 2001
treatment: cnwb_pooled, wb_only, cn_only

The recovered pair outputs follow a pattern similar to:

./data/matches/matches11_{treat_type}_{region}{level}T{T}{y0}.csv

Matched trios were requested conceptually, but the recovered implementation should be treated as unconfirmed until the actual trio output files are found and validated.


13. Matching-specific diagnostics

The following diagnostics remain necessary whenever matching is selected as the estimator family. They should now be understood as a subset of the broader experiment-validity gate sequence.

13.1 Minimum diagnostic checklist

DiagnosticQuestion
Treated NHow many treated area-periods exist for this design?
Control NHow many candidate pure-control area-periods exist?
Matched NHow many treated units were successfully matched?
Outcome coverageHow many matched units have valid post-treatment outcome data?
Covariate completenessHow many units are dropped because of missing covariates?
Balance improvementDo treated/control covariate differences shrink after matching?
Common supportAre treated units comparable to available controls?
Source sparsityDoes one source family, especially China, sharply limit sample size?
Geographic plausibilityAre matched units plausible comparisons within the same broad geography and time period?
SensitivityDo results change across ADM level, time window, and treatment definition?

13.2 Known matching risks

  • ADM1 areas may be too large and heterogeneous for credible matching.
  • Chinese project exposure appears sparser than World Bank exposure, especially after intersecting with outcome data.
  • Multi-location projects may dominate exposure counts if project-location weighting is not handled carefully.
  • Project classification uncertainty propagates directly into treatment uncertainty.
  • Missing covariates may restrict the sample in non-random ways.
  • Outcome coverage, especially for survey outcomes, may be much thinner than exposure coverage.
  • Matching on observed covariates does not by itself solve endogenous project placement on unobserved factors.

14. Estimator layer

Regression notebooks exist, but they should be treated as prototypes inside a broader estimator family until canonical experiment specifications are agreed.

14.1 Candidate estimator families

FamilyDescriptionStatus
Raw / descriptive comparisonCompare treatment groups, pre/post distributions, trajectories, and support before causal interpretation.Required diagnostic baseline.
Matching-based comparisonCompare post-treatment outcomes between treated and matched control units; may include matched-sample regression.Recovered implementation exists.
Completed vs planned spatial comparisonCompare exposure around implemented/completed projects with future/planned project locations selected through a similar siting process.Methodologically relevant; FCV timing/status feasibility must be validated.
Longitudinal / staggered-treatment designUse repeated area-time outcomes around implementation timing and compare treatment cohorts over event time.Promising for ACLED/UCDP; requires modern staggered-treatment estimators rather than naïve TWFE by default.
Count / rate modelPoisson, negative-binomial, binary, hurdle, or rate models for sparse violence outcomes.Outcome-model choice remains open.
Multi-arm comparisonJobs vs non-jobs vs pure control or WB vs China vs pure control.Requires stable labels and sufficient support.
Spatial spillover-aware designExplicitly model treatment rings, neighboring exposure, or interference rather than assuming nominal controls are uncontaminated.Future extension after the baseline apparatus is validated.

14.2 Estimator descriptors to freeze

Before reporting results, the team should freeze these descriptors:

DescriptorNeeded decision
Outcome transformationRaw count, binary indicator, log transform, per-capita rate, or category-specific outcome.
Treatment timingTreatment window, project status convention, and post-treatment outcome window.
CounterfactualPure control, non-jobs investment, future/planned project locations, not-yet-treated areas, or another explicit comparison group.
Covariate timingWhich covariates are pre-treatment and how they are lagged.
Geography / exposureADM level, project-to-area assignment, radius, or other spatial exposure rule.
Fixed effectsNone, time, country, area, matched-pair, or combinations appropriate to the estimator.
Clustering / uncertaintyArea, survey community, country, pair, bootstrap, spatially robust, or other justified errors.
SampleAll area-periods, matched pairs only, planned/completed sample, event-study cohorts, multi-arm subsets, or source-specific subsets.
Treatment definitionAny investment, WB-only, CN-only, jobs-any, direct jobs, indirect jobs, non-jobs.
InterpretationDescriptive association, matched comparison, spatial status contrast, panel estimate, or causal estimate under explicit assumptions.
Robustness gridAdmin level, time window, y0, source family, outcome family, classification rule, bandwidth, and estimator family.

14.3 Cross-estimator validity and falsification gates

Regardless of estimator family, the analysis should expose the following before a substantive claim is made:

GateQuestion
Data integrityAre keys, dates, duplicates, joins, source fields, and outcome definitions valid?
Timing resolutionCan project state and treatment timing be assigned without guessing?
Treatment discriminabilityDo treatment categories have substantive and empirical contrast, or does a broad category collapse almost everything into treatment?
Support / overlapHow many observations actually identify the contrast after filters, fixed effects, clustering, and missingness?
Spatial precisionIs the requested exposure resolution compatible with source geocoding uncertainty?
Exposure collisionHow often do multiple donors, projects, sectors, or temporal states overlap?
Selection diagnosticHow do future/planned or treated groups differ from never-treated controls before treatment?
Outcome sensitivityIs the outcome too sparse/noisy for the proposed estimand?
Synthetic signal recoveryIf a plausible effect were injected, how often would the experiment recover it?
Placebo / falsificationDoes the design generate effects on pre-treatment outcomes, fake timing, negative controls, or implausible exposure definitions?
SensitivityDoes the sign or magnitude depend entirely on one arbitrary ADM level, window, radius, or coding rule?

These gates are designed to make failed experiments informative. They do not guarantee that a real substantive effect will be found.


15. Decision table for current candidate experiments

This table should guide the next discussion with Eric and Charlotte.

DesignData exists?Labels needed?Outcome coverageSample-size riskReady for estimation?Recommended next action
WB any investment vs pure controlYes-ishNoACLED likely strongestMediumMaybe, after gatesValidate panel, treatment timing, support, outcome coverage, and placebo behavior.
CN any investment vs pure controlYes-ishNoACLED likely but sparseHighNot yetQuantify treated N and outcome coverage by ADM/T/y0.
WB+CN pooled vs pure controlYes-ishNoACLED likelyMediumMaybe, after gatesUse as broad calibration experiment, not final theory test.
WB jobs-any vs pure controlPartialYesACLED/Afrobarometer TBDHighNoComplete annotation protocol and labels.
CN jobs-any vs pure controlPartialYesSparse / TBDVery highNoCheck updated China/AidData source and sample viability.
Jobs-any vs non-jobs vs pure controlConceptual / partialYesLikely sparseVery highNoLocate or rebuild multi-arm design after labels.
Direct jobs vs indirect jobs vs pure controlNot readyYesTBDHighNoUse only after broad jobs-any labels are stable.
Completed vs planned project exposureSource-field dependentMaybeOutcome-specificUnknownNot yetAudit agreement/start/end/status fields and construct support/composition diagnostics.
ACLED longitudinal / event-time designPanel structure existsNo for source-family treatmentStrongest temporal coverageMediumNot yetValidate timing and support; choose a modern staggered-treatment estimator family.
Afrobarometer spatial designData work likely partialMaybeTBDHighNoProduce coverage table by country, survey round, project status, exposure radius, and treatment.
DHS outcomesUnclearMaybeTBDMedium/highNoDecide whether DHS is an outcome source or covariate source.

16. Open decisions

The following decisions should be made before the pipeline is presented as final.

16.1 Source and version decisions

  1. Which World Bank public project dataset is canonical?
  2. Which AidData World Bank geocoded release is canonical?
  3. Which Chinese development finance dataset is canonical?
  4. Is there a newer source that supersedes the 2023-era AidData/TUFF files?
  5. Should old analysis be replicated first, or rebuilt directly on updated data?
  6. Which source fields can support agreement, planned, start, implementation, and completion timing?

16.2 Annotation decisions

  1. What is the first-stage label: jobs_any vs non_jobs, or direct/indirect from the start?
  2. How should mixed projects be classified?
  3. How should macro-policy-only projects be excluded or flagged?
  4. What confidence threshold is required before labels enter an experiment?
  5. Should labels be human-coded, rule-based, ML-assisted, or a hybrid?

16.3 Treatment and counterfactual decisions

  1. Should exposure be binary or intensity-based?
  2. If intensity-based, how should multi-location project amounts be split?
  3. Should treatment be based on approval date, commitment date, start date, implementation period, completion date, or disbursement timing?
  4. Can future/planned project locations be constructed credibly enough to serve as a counterfactual family?
  5. How should overlapping WB and China projects be treated?
  6. How should repeated exposure across time windows be handled?
  7. How should ambiguous project status be represented rather than silently guessed?

16.4 Geography, support, and estimator decisions

  1. What is the main administrative level: ADM2 or ADM3?
  2. What is the main time window: 2, 3, or 4 years?
  3. For respondent/project linkage, what radius is substantively plausible and compatible with geocoding precision?
  4. Are matched trios necessary for the first analysis, or should pairs come first?
  5. What is the minimum acceptable effective treated N and outcome N after all design restrictions?
  6. Which estimator family is the first canonical calibration design?
  7. Which placebo and synthetic-signal tests must pass before interpreting the first real coefficient?

17. Proposed next implementation architecture

The recovered notebooks should not be discarded. They should be converted into a modular pipeline in small steps. The newer fcv-experiment-harness provides an initial validation/experiment layer, while the structure below remains a useful map for future source ingestion and canonical-data modules.

17.1 Proposed modules

fcv_investments/
config/
sources.yaml
geography.yaml
time_windows.yaml
treatment_definitions.yaml
outcomes.yaml

data_registry/
raw_sources_manifest.csv
derived_outputs_manifest.csv
validation_report.md

src/
ingest/
wb_public.py
wb_aiddata.py
china_aiddata.py

spatial/
point_to_gid.py
aggregate_area_time.py
exposure_intensity.py

classify/
label_schema.py
apply_project_labels.py
build_annotation_table.py

panels/
build_area_period_panel.py
merge_covariates.py
merge_outcomes.py

matching/
build_pairs.py
build_trios.py
diagnose_balance.py
sample_coverage.py

regressions/
run_baseline.py
run_robustness_grid.py
summarize_results.py

The active experiment harness should sit downstream of canonical data construction rather than duplicate source-ingestion logic.

17.2 Proposed notebooks after migration

notebooks/
01_source_inventory.ipynb
02_treatment_construction_validation.ipynb
03_annotation_review.ipynb
04_sample_viability_tables.ipynb
05_matching_diagnostics.ipynb
06_regression_results.ipynb

The notebooks should become review surfaces. Production logic should move into reusable functions.


18. Immediate next artifacts

To make this project usable for Charlotte and reviewable by Eric, the next deliverables should be:

  1. Annotation Protocol Memo
    Defines project-level labels and coding rules.

  2. Experiment Manifest Table
    Defines candidate experiments by treatment, timing, geography/exposure, counterfactual, outcome, sample, estimator family, and status.

  3. Source-Version Crosswalk
    Compares WB public data, AidData WB geocoded data, and China/AidData data by unit, ID, year coverage, timing/status fields, location coverage, and use.

  4. Sample-Viability / Gate Table
    Reports treated N, control/counterfactual N, outcome coverage, admin level or radius, time window, treatment definition, placebo status, signal-recovery status, and readiness.

  5. Minimal Canonical Experiment Spec
    One baseline design only, initially chosen for measurability rather than because it produces a favorable coefficient.

The Validation Status page should remain the human-facing summary of which real-data experiment surfaces have actually passed these checks.


The next meeting should not try to review every notebook. It should answer this question:

Which empirical experiment should be treated as the first canonical analysis path, given source validity, treatment timing, counterfactual credibility, effective sample support, outcome coverage, and the initial gate results?

A good meeting outcome would be one of:

  • Choose WB-only before CN because sample coverage is stronger.
  • Choose jobs_any first, postponing direct/indirect distinctions.
  • Use ACLED first, postponing Afrobarometer until coverage is documented.
  • Treat ADM2 as the main area-period unit and ADM3 as robustness.
  • Test whether planned/completed project status is feasible before relying only on pure-control matching.
  • Rebuild source data using updated AidData before producing final labels.
  • Produce the sample-viability / gate table before interpreting regression reruns.

20. Bottom line

The recovered notebooks show that the empirical pipeline was substantially implemented, but not consolidated into a stable research software architecture or a fully explicit causal design.

The recovered scientific spine is:

project finance -> geocoded exposure -> area-period treatment -> matched comparisons -> post-treatment outcomes

The active scientific spine is now broader:

canonical empirical data
-> experiment specification
-> analysis sample
-> validity / measurement gates
-> estimator family
-> falsification and sensitivity
-> research interpretation

The unresolved bottlenecks are:

  1. source-version and project-timing validation;
  2. project annotation and treatment definition;
  3. counterfactual choice, including whether planned/future project locations are feasible;
  4. effective sample support and outcome coverage;
  5. geolocation precision and exposure collisions;
  6. placebo/falsification behavior;
  7. empirical sensitivity to plausible weak effects;
  8. selection of one first canonical estimator family after the preceding gates are visible.

The safest next step is not to rerun everything or to search across regressions for a favorable result. The safest next step is to connect one recovered real-data experiment to the validation harness, observe which gates fail, repair or narrow the design where necessary, and only then interpret the estimator output.