SIMTELLIGENCE
← All field notes
Continuity / Research essayAugust 2026 · 3 min read

The next model will arrive. What should survive?

Every model upgrade offers a fresh possibility. A software factory needs a way to tell whether it is actually an improvement.

An interchangeable teal engine suspended above an ivory machine with alternative modules nearby
A replaceable model needs a stable system around it.

A new model arrives with an impressive demonstration. It handles a task that seemed difficult a month ago. The obvious question is whether to switch.

For an autonomous software factory, there is a second question: what else would have to change? If adopting a model means rewriting every workflow, replacing the checks, and relearning how costs behave, the upgrade reaches far beyond the assistant.

The research problem is continuity. A factory needs to improve as models become more capable while retaining a dependable way to organise work and judge results.

Keep the measuring instrument steady

Suppose a new model completes a maintenance task faster. That sounds promising. But perhaps the tool access changed at the same time, the budget increased, or the test became easier. The result describes a different system, making the model’s contribution harder to identify.

A useful comparison holds the surrounding conditions steady where possible: the task, available tools, permissions, time limits, and grading criteria. Necessary differences should be recorded so readers can understand the limits of the comparison.

This is less dramatic than swapping the model and admiring the answer. It is also more useful for deciding whether to entrust the new system with real work.

The assets between releases

Several parts of a factory should remain valuable through an upgrade. A clear tool interface still explains what an action does. A reliable test still catches the defect it was designed to catch. A release record still tells the next operator what was shipped.

The same applies to lessons from failures. If a run lost track of a changed requirement, the useful finding concerns how requirements are recorded and revisited. That finding can inform the next system even when the model changes.

Some instructions will need revision. Some tools will become unnecessary. Continuity does not mean freezing the design. It means understanding which changes are improvements and which are simply adjustments to a new component.

An upgrade is easier to judge when the factory remembers how the previous version performed.

Make adoption an experiment

A practical proposal is to begin with a limited set of representative tasks. Run the current system and the candidate under documented conditions. Compare correctness, review effort, time, cost, and the kinds of failure each produces.

A mixed result is useful. A model might be better at planning and less dependable at following a particular tool contract. The appropriate response could be narrower adoption, a change to the interface, or further testing.

This makes model selection an ongoing operational decision. The factory can adopt new capabilities without treating every announcement as a reason to rebuild itself.

The enduring research is in that ability to learn: a record of what changed, a method for testing the next claim, and a system that remains understandable while its components evolve.

Continue the inquiry

What would the next model comparison teach us that remains useful after another upgrade?

Send a perspective ↗

Read next / Architecture

The machine around the machine