What I wanted
I write essays and notes about agents, and I wanted something to attack them before anyone else does: find the claims, the jumps between them, and the parts I only hope are true. Then research the weak spots and tell me what to do next. On September 22 I asked Codex to turn that into a Method, to "root out any wishful thinking, nonsense, illogical jumps".
Version 1 had five agents
The first version was 733 lines of YAML with five agent steps: define the claim, plan the research, collect evidence, assess the claim, and save. Every step had a strict evidence check. My first questions were "is this 5 separate codex agent calls?" and then "that seems extremely excessive."
Three steps were enough
By version 4 it had three steps: one model call plans the research, one agent does the research, and a script checks and saves the result. I liked how it broke each claim into a fixed shape (Claim, Depends on, Finding, Consequence), so that stayed. The report was too long and dense, so version 7 removed the repeated text.
Stricter checks failed good runs
Versions 9 to 11 added integrity checks, and only 3 of 11 runs passed. Some failures were real, such as a reasoning link that was reworded without a record. Others rejected correct work: one check refused evidence that the input itself supplied. Version 12 fixed the instructions and loosened that check. It passed 8 of 9 runs, including five test cases: ambiguous notes, a causal claim, a comparison of two pass rates, a values claim, and a claim that mixed up training and evaluation.
| Versions | Runs passed | What changed |
|---|---|---|
| 1 to 8 | 7 of 12 | Five agents became three steps; shorter report |
| 9 to 11 | 3 of 11 | Stricter integrity checks |
| 12 | 8 of 9 | Fixed instructions, looser check, five test cases |
| 13 and 14 | 5 of 8 | Version 13 accepts citations to the supplied sources |
| 15 to 17 | 0 of 3 | New steps and eight new checks |
| 18 | 1 of 1 | Same steps, seven checks |
It was not searching the web
On October 2 I ran version 14 on a claim from one of my product ideas. It returned a tidy map of my argument, but it had not looked for anyone else's work on the idea. I wrote back: "the whole point of this method is to RESEARCH a vague claim and find evidence for it."
Eight new checks broke it again
Version 15 added the missing research: questions about who else has tried the idea, a log of every search, a plan that I approve before research starts, a separate assessment, and a step that argues against the conclusion. It also added eight new checks. Version 15 timed out. Version 16 failed after 14.8 minutes because one opened page had no row in the search log. I stopped version 17 and wrote that this "sounds like a bunch of brittle deterministic tests that just fail things unnecessarily."
Version 18 kept seven checks
Version 18 kept the new steps and dropped every check that did not prevent a specific harm. Everything else is listed in the report instead of failing the run.
| Step | Check | What it prevents |
|---|---|---|
| plan | The claim is unchanged | My claim being rewritten |
| plan | My decision is unchanged | My decision being replaced |
| plan, assess | Every quote attributed to me is in my input | Words I never wrote |
| research | Every finding cites a source that exists and was checked | Invented evidence |
| research | Every outside source has a real URL or path | Sources nobody can check |
| assess | Every verdict cites recorded evidence | Verdicts without evidence |
| save | The saved files equal the accepted record | A corrupted result |
What changed on the same claim
I ran versions 14 and 18 on the same claim and the same decision.
| Version 14 | Version 18 | |
|---|---|---|
| Outside sources | 31 | 59 |
| Findings | 21 | 30 |
| Searches and opened pages logged | 0 | 130 |
| Statements in the argument map | 39 | 29 |
| Run time | 12.4 min | 19.1 min |
| Report | 48 KB | 76 KB |
Version 18 is slower and its report is longer, mostly because it now lists the prior work it found. The argument map got smaller, because it stopped splitting one claim into several.
What I learned
A check should prevent a harm I can name. When a check fails correct work, the run throws away minutes of research and tells me nothing new. Twice, removing checks fixed this Method. The authoring guide in the Method CLI says the same: no extra check unless it catches a failure that matters.
The current shape
The full file has the prompts and output types. This is its outline:
steps:
plan: call # reconstruct the argument, plan the research (checked)
present_plan: run # show me the plan
confirm: ask # I approve or change the plan
research: agent # search, read, and log every source (checked)
assess: call # judge each claim from recorded evidence only (checked)
challenge: call # the strongest objection to the conclusion
save: run # write the report and the record (checked)
