← All posts

How I built a claim research agent

I wanted a tool that reads my own essays, maps the claims, and researches the weak ones before I publish. This is how the Method changed over 18 versions and 44 runs.

What I wanted

I write essays and notes about agents, and I wanted something to attack them before anyone else does: find the claims, the jumps between them, and the parts I only hope are true. Then research the weak spots and tell me what to do next. On September 22 I asked Codex to turn that into a Method, to "root out any wishful thinking, nonsense, illogical jumps".

Version 1 had five agents

The first version was 733 lines of YAML with five agent steps: define the claim, plan the research, collect evidence, assess the claim, and save. Every step had a strict evidence check. My first questions were "is this 5 separate codex agent calls?" and then "that seems extremely excessive."

Three steps were enough

By version 4 it had three steps: one model call plans the research, one agent does the research, and a script checks and saves the result. I liked how it broke each claim into a fixed shape (Claim, Depends on, Finding, Consequence), so that stayed. The report was too long and dense, so version 7 removed the repeated text.

Stricter checks failed good runs

Versions 9 to 11 added integrity checks, and only 3 of 11 runs passed. Some failures were real, such as a reasoning link that was reworded without a record. Others rejected correct work: one check refused evidence that the input itself supplied. Version 12 fixed the instructions and loosened that check. It passed 8 of 9 runs, including five test cases: ambiguous notes, a causal claim, a comparison of two pass rates, a values claim, and a claim that mixed up training and evaluation.

VersionsRuns passedWhat changed
1 to 87 of 12Five agents became three steps; shorter report
9 to 113 of 11Stricter integrity checks
128 of 9Fixed instructions, looser check, five test cases
13 and 145 of 8Version 13 accepts citations to the supplied sources
15 to 170 of 3New steps and eight new checks
181 of 1Same steps, seven checks

It was not searching the web

On October 2 I ran version 14 on a claim from one of my product ideas. It returned a tidy map of my argument, but it had not looked for anyone else's work on the idea. I wrote back: "the whole point of this method is to RESEARCH a vague claim and find evidence for it."

Eight new checks broke it again

Version 15 added the missing research: questions about who else has tried the idea, a log of every search, a plan that I approve before research starts, a separate assessment, and a step that argues against the conclusion. It also added eight new checks. Version 15 timed out. Version 16 failed after 14.8 minutes because one opened page had no row in the search log. I stopped version 17 and wrote that this "sounds like a bunch of brittle deterministic tests that just fail things unnecessarily."

Version 18 kept seven checks

Version 18 kept the new steps and dropped every check that did not prevent a specific harm. Everything else is listed in the report instead of failing the run.

StepCheckWhat it prevents
planThe claim is unchangedMy claim being rewritten
planMy decision is unchangedMy decision being replaced
plan, assessEvery quote attributed to me is in my inputWords I never wrote
researchEvery finding cites a source that exists and was checkedInvented evidence
researchEvery outside source has a real URL or pathSources nobody can check
assessEvery verdict cites recorded evidenceVerdicts without evidence
saveThe saved files equal the accepted recordA corrupted result

What changed on the same claim

I ran versions 14 and 18 on the same claim and the same decision.

Version 14Version 18
Outside sources3159
Findings2130
Searches and opened pages logged0130
Statements in the argument map3929
Run time12.4 min19.1 min
Report48 KB76 KB

Version 18 is slower and its report is longer, mostly because it now lists the prior work it found. The argument map got smaller, because it stopped splitting one claim into several.

What I learned

A check should prevent a harm I can name. When a check fails correct work, the run throws away minutes of research and tells me nothing new. Twice, removing checks fixed this Method. The authoring guide in the Method CLI says the same: no extra check unless it catches a failure that matters.

The current shape

The full file has the prompts and output types. This is its outline:

steps:
  plan:          call   # reconstruct the argument, plan the research   (checked)
  present_plan:  run    # show me the plan
  confirm:       ask    # I approve or change the plan
  research:      agent  # search, read, and log every source            (checked)
  assess:        call   # judge each claim from recorded evidence only  (checked)
  challenge:     call   # the strongest objection to the conclusion
  save:          run    # write the report and the record                (checked)

Make the next run better.