Agents: read llms.txt

Simon Spoorendonk

1 October 2026

How I check agent-built solvers

Nobody reads all the code any more, and I don’t either. What I care about is keeping an agent from handing me an algorithm, a solver or a paper that looks right and isn’t.

My examples are from building optimization algorithms. I think most of it carries over to other code that implements a published method and produces numbers.

This is how I work as of autumn 2026. I started building with coding agents in February 2026, and it took off properly in April. I focused on a specific kind of project: one or a few papers turned into a new code base, run through experiments, and written up. Greenfield code, no users to keep happy, and not necessarily novel research. Some of it is reimplementation, some is engineering, and the paper should say which is which. mip-heuristics, for example, puts four published heuristics inside one solver so they can be compared like for like. That’s a benchmark, not a new method.

A flow chart. Along the top, the source papers lead to benchmarks, then epics and issues, then code. Down the right, code leads to a code review against the sources, then along the bottom to checks, a paper review against the code and the citations, and finally my paper. Blue arrows run from the code review and from the checks back to the code. A green arrow runs from the checks back to the source papers.
How a project runs, from other people's papers to mine. Red is the main flow. Blue arrows go back to the code when a review or a check fails. The green arrow goes back to the source papers, when the code doesn't behave the way a paper says it should. Before my paper goes out, it's reviewed against the code and its citations against the papers they cite.

Reading the papers

I don’t read every paper either, and certainly not when the trail goes back twenty years. The agent reads them. What I do is cross-examine it. What does this paper do, how is it different from the one before, is that in the paper or did you infer it. I keep asking until the answers stop moving. Having read a lot of papers in my earlier days helps with knowing what to ask.

It doesn’t catch everything. Discrepancies still turn up later, usually when something is implemented and doesn’t behave the way the paper says it should.

What do we compare against?

Before planning any code I decide how I’ll validate it. The best case is an answer to compare against, and I prefer answers that are published or that I can reproduce, such as known optima on public instances. Results from old papers work too, the solution values rather than the run times, since the machines are long gone. So does other people’s published code, if it runs.

When there’s no reference solution at all, the ablations have to carry more of the weight. An ablation switches components off one at a time to see what each one is worth. Ablations can’t tell you the answer is right, but they can tell you sensible things about each part of the algorithm and whether it does what it’s supposed to.

Epics, issues and docs

I track the work in GitHub epics: development, experiments, and writing the paper. The experiments split in two. The headline runs are usually given up front, since the point is speed or solution quality on benchmarks everyone already uses. Ablations are harder to plan, and they drift as I understand the code better. That’s fine as long as the drift is written down.

Under the epics, the agents write issues to themselves. I rarely read them, but we talk about them, and they’re the project’s memory. Anything decided, found or postponed ends up in one.

The docs in the repo take more effort than you’d think. Agents write a lot of documentation, and most of it is too long. Keeping it short and to the point is work I have to ask for, again and again.

Did we build the algorithm in the paper?

Most of the review rounds go on this. The point is to make sure the code implements the algorithm the paper states, all of it, and that nothing was left out because the agent postponed it or decided it wasn’t needed. The code should only differ from the paper where the paper is unclear, where we found a better way, or where the authors’ own code does something different and we follow it. Either way it gets written down, and a change like that usually needs its own tests.

On mip-heuristics the review found many defects across the four heuristics it implements. In one, the relaxed solution the method is built around never reached the rounding step. In another, what shipped was the variant the paper itself abandons. It also turned up two bugs in Feasibility Jump inside HiGHS. Both came in with the reference implementation. The objective was counted with the wrong sign, so it steered towards worse solutions, and when working out where to move a variable, rows where it has a negative coefficient were skipped.

Nothing in the benchmark results pointed at any of them. A heuristic running the wrong variant still runs, and still finds solutions some of the time. Reading the code against the paper is what found them.

Alongside that there are regular reviews of the architecture and the performance, like any project should have.

Checks on the numbers

For bugs where the answer comes out wrong, the checks are the usual ones. Tiny instances you can enumerate. The reference answers from before any code existed. The same problem through someone else’s solver on the same machine. Unit tests catch the mechanical stuff. Ablations are about which component carries the result, but because they mean running a lot, they turn up the odd bug as well.

Bugs in my own logic are cheap. These checks find them, and they’re quick to fix because the whole code base is right there. The expensive ones are in code I didn’t write. In mcfcg, NVIDIA’s GPU solver cuOpt once handed back a cost about 170 times too high, sometimes labelled solved, after a linear algebra step failed underneath it. I fixed that one in my own fork. HiGHS’s new interior-point solver threw away a near-optimal solution and reported an internal error. In both cases it took another solver on the same model to see what had gone wrong.

Writing the paper

This is where a lot of my own time goes.

The algorithms in the paper are written the way they’re implemented, and where they’ve changed from a previously published version, the paper says so.

Citations get checked twice. The agent follows each DOI to check the reference exists and says what we think it says, and then there’s whether the person cited did what we cite them for. I’m nitpicky about that one, especially in the literature overview. Giving the right people credit matters.

No number is typed from memory. In the multi-commodity flow paper the tables are generated by script from the result files. A second script holds a copy of every number in the prose that comes from our runs, re-derives each one from the data, and fails if they disagree. The checker has had bugs too. It once let a 1.94 through where the table said 1.93.

And I try hard not to overclaim. A result goes in with the conditions it holds under and where it loses.

Same as always, just faster

None of this is new. It’s the project, software and research practice that good work always needed: decide how you’ll validate, check the code does what the paper says, don’t trust a number nobody derived, and credit people properly.

What’s changed is the speed. The work goes a lot faster, but these algorithms are still hard to implement, and an agent will hand you something that looks finished long before it is.

That’s where things are now.