Introducing ReplicateAI

Point an agent at an applied-econ PDF and a public CSV, estimate the headline coefficient in a Modal sandbox, and get a referee-style audit back.

Most agent demos succeed by synthesizing prose. Empirical replication is harder in a useful way: the published β does not care how confident the model sounds, and almost every real run fails at least once on a dtype, a missing column, or a bad specification.

I wanted a narrow demo of that loop. Paper PDF plus public CSV in, a referee-style audit out, with the debugging visible. That project is ReplicateAI. It is a portfolio research tool with curated example packs, not a “replicate any paper” product.

What a run does

Point the CLI at an example pack (or a folder with paper.pdf and data.csv):

cd replicate_ai
uv sync
uv run replicate-ai ../examples/card_krueger

# CI / plain stdout
uv run replicate-ai --no-tui ../examples/card_krueger

# Browser launcher (uv sync --group gui first)
uv run replicate-ai --gui

Host-side PDF preflight (Docling by default) writes paper_text.md and paper_tables.json. Those files, plus the PDF and CSV, go into a Modal /workspace. A LangChain Deep Agent acts as econometrician: it locks a target_specification.json, writes estimation scripts, runs them in a Python sandbox (statsmodels, linearmodels, and friends), and edits when the traceback says so. An auditor sub-agent compares the estimate to the published number and writes replication_audit.md with MATCH, CLOSE, MISMATCH, or FAILED.

On a TTY you get a Textual dashboard with phases, a live log, and a headline card. --gui opens a browser launcher. --no-tui is the path for CI.

Why Deep Agents and Modal

You should expect wrong code, long logs, and a success criterion that is just a number you can check against a table. Deep Agents already has the harness pieces that make that tractable: planning that survives context flushes, a virtual filesystem that spills fat stdout, and sub-agents with their own windows. PDF extraction stays on the host so the sandbox image stays lean. Econometrics stays in Modal so a broken script cannot trash the laptop.

The auditor is strict on purpose. MATCH means same sign, relative deviation within 5%, and the same significance bucket. That is closer to how a referee reads a table than how a chatbot grades itself.

Six packs, and a MATCH on the wrong estimand

Six packs under examples/: Card & Krueger, Dehejia-Wahba, Imbens-Rubin-Sacerdote, Angrist-Lavy, Autor-Dorn-Hanson, and Acemoglu-Johnson-Robinson. Public data only. One headline estimand per run.

Documented Anthropic runs in docs/test.md:

PackVerdictNotes
Card & KruegerMATCHOne run hit 2000-reply coeffs, not 1994 Table 3
Dehejia-Wahba (NSW experimental)MATCHHeadline treat → RE78
Imbens lotteryMATCHPrize-on-earnings elasticity, ~4% rel. dev.
Autor-Dorn-HansonMATCHImport exposure → mfg share
Angrist-LavyCLOSEClass-size IV, ~7% off
AJR(empty)Still unfilled

The MATCH count is less interesting than the failure modes. Take Imbens. An earlier run MATCH’d a different table than the pack documents: numerically tight on the coefficient the agent chose, wrong for the pack target. The auditor did its job on the number it was handed. The miss was upstream, in which estimand got locked.

Wrong-estimand MATCH is the failure mode I care about most, because it looks like success. A prose agent can hallucinate a coefficient and you might not notice until a human opens the PDF. Here the number is real, the regression ran, the relative deviation cleared 5%, and you still replicated the wrong claim. That is a worse bug than MISMATCH. MISMATCH at least tells you to keep digging. A green MATCH on the wrong table makes it easy to stop looking.

Card & Krueger has a related miss in the notes: one run hit 2000-reply coeffs instead of 1994 Table 3. Same paper family, different published number, still a MATCH if you score the wrong target. So the roadmap starts with a BENCHMARK.md (pack × provider × verdict × failure tag) instead of more papers. I would rather know how often the agent picks the wrong estimand than how many packs I can claim as green.

What I would not claim

The sandbox is Python only: there is no Stata or R path, no credentialed microdata, and no full-table replication. This is also not production hosting. The README says that out loud because the claim worth making is autonomy on a narrow, checkable task, not coverage of the AER archive.

If you want to try it, start with Card & Krueger (PDF and data ship in-repo), then work through the other packs in examples/. Issues and design notes live in the repo.