In early March 2026, Andrej Karpathy released a project called autoresearch. The premise was simple: give an AI agent a training script, a metric, and a prompt that says "make this better." Let it run.
Within days, community members were reporting results: 126 experiments over ~10.5 hours, driving validation loss from 0.9979 to 0.9697. The agent hypothesized improvements, modified the code, measured the results, kept what worked, discarded what didn't, and repeated — roughly twelve times per hour. Separate runs showed that the best improvements transferred cleanly to larger models.
Karpathy still set the objective, the constraints, and the proxy setup. He still manually verified transfer to larger models. But the research loop itself — the part that used to require a grad student staring at loss curves for weeks — ran autonomously.
That shift should sit with you for a moment, because Karpathy built this for ML training. But the loop he formalized — hypothesize, modify, measure, keep or discard — has nothing to do with machine learning specifically. It applies to anything you can measure.
Test suite runtime. Endpoint latency. Page load speed. Bundle size. Memory consumption. P99 response time. Throughput.
If you can measure it, you can optimize it autonomously.
This is not a theoretical observation. Within days of the release, Shopify CEO Tobi Lütke ran autoresearch against Liquid — Shopify's templating engine, powering millions of storefronts. The result: 53% faster combined parse+render time, 61% fewer object allocations. "This is probably somewhat overfit," he wrote, "but there are absolutely amazing ideas in this."
That's a CEO pointing an optimization loop at core infrastructure and getting a 53% speedup overnight. Not a research demo. Production code.
I've been running these same loops against my own codebases, and the results are changing how I think about software.
The Numbers
Here's what I've seen running autonomous optimization loops across several codebases I work with. These are my own results — not published benchmarks — but I'm sharing them because the pattern is more interesting than any single number:
- Test suites: 30–40% runtime reductions. Not by deleting tests — by restructuring setup, parallelizing where possible, eliminating redundant database operations. Across multiple codebases, consistently.
- Endpoint latency: An API endpoint that was already serving at ~23.5μs average — fast enough that most engineers wouldn't bother optimizing it further. The loop drove it to ~5.0μs. That's 4.7x faster. P99 dropped from ~35μs to ~13μs. Throughput went from ~70K ops/sec to ~200K ops/sec. This is what's interesting: the endpoint was already well-optimized by conventional standards. The loop found headroom that a human would have walked past.
- Page loads: 2-second improvements on frontend render paths. Not heroic rewrites — incremental changes that compound.
- JVM tuning: Heap sizes, garbage collection parameters, object allocation patterns — the kind of configuration that most teams set once and never revisit because the parameter space is too large to explore manually. An optimization loop can try hundreds of combinations against your actual workload and find configurations that meaningfully reduce pause times and memory pressure.
None of these required a performance engineer. None required someone to sit down and profile the code. An agent ran the loop: try something, measure, keep or discard. Hundreds of iterations. The compounding effect is what matters — each improvement is small, but the aggregate is not.
The pattern is deceptively simple: (1) Agent reads the code and a set of human-written
instructions. (2) Agent forms a hypothesis. (3) Agent modifies the code. (4) The system measures
against a fixed metric. (5) Agent keeps improvements, discards failures. (6) Repeat. Karpathy's
system ran ~12 experiments per hour against a fixed 5-minute training budget. The same loop works
for any metric with a fast feedback cycle.
These aren't cherry-picked wins from exotic systems. They're routine improvements on ordinary codebases. The kind of optimizations that senior engineers would get to "eventually" — except eventually never comes because there's always a feature to ship.
The loop doesn't have a backlog. It doesn't get distracted. It runs while you sleep.
The Organizational Shift
Here's where it gets interesting. If you have an autonomous optimization loop running continuously against your codebase, the calculus for how you build software changes.
The traditional sequence: build the feature, write clean code, optimize where it matters, ship when it's polished. Every step requires the same scarce resource — a skilled engineer's attention.
The new sequence: build the feature, ensure it works, ensure it has tests, ship it.
That's not a joke. If you have loops that can autonomously find and eliminate performance waste, redundant computation, and suboptimal patterns — while verifying that tests still pass — then agonizing over internal code quality at write-time is misallocated effort.
Let your product engineers go wild. Let them build the feature with whatever tools and patterns get them to a working, tested implementation fastest. Don't worry as much about whether the code is beautiful. Worry about whether the tests are comprehensive and the acceptance criteria are clear.
Then let the loop run.
The slop gets optimized away. Not by a human painstakingly refactoring each function, but by an agent that tries a hundred variations and keeps the ones that are faster, smaller, or more efficient — with tests passing at every step.
This inverts the usual engineering priority stack:
1. Test coverage becomes the highest-leverage investment. Not because tests catch bugs (though they do), but because tests are the contract that lets autonomous optimization happen safely. If you can't measure it, you can't improve it. If you can't verify correctness, you can't accept changes.
2. Acceptance criteria become the product. Specification quality determines what the loop can optimize _toward_.
3. Internal code quality becomes a downstream consequence, not an upstream concern. The code gets better as a side effect of optimization, not as a goal in itself.
This is a significant reordering. It says: invest in the _boundary conditions_ — what the code must do, how fast it must do it, what correctness looks like — and let the interior be machine-managed.
The Quality Tension
I'm aware this cuts against something I've argued before. In Quality Is the Bottleneck Now, I made the case that quality lives in feedback loops, not in code internals. That "internals matter less than they used to." That evidence-based verification is the real work.
Autoresearch is the logical endpoint of that argument.
If quality is a property of the feedback loops around the code — not the code itself — then the ultimate feedback loop is one that runs autonomously, measures continuously, and improves without human intervention.
But it also sharpens the distinction between layers of quality:
- Behavioral quality — does the software do what it should? — still requires human judgment. Acceptance criteria, UX coherence, product intent. No loop optimizes for "does this feel right."
- Measurable quality — is the software fast, efficient, reliable? — is exactly what autonomous loops excel at. Give it a metric, give it tests, let it work.
- Internal quality — is the code readable, well-structured, elegant? — becomes a different kind of concern entirely. Not unimportant. But increasingly a property that exists for machines to manage, not humans to enforce.
I want to sit with that tension rather than resolve it cleanly, because I think the honest answer is: we're in transition. We still read code. We still review diffs. But the pressure is clearly toward a world where the interior of our software is optimized by processes we designed but didn't execute — and the code that results isn't written for human aesthetics.
The Review Problem, Compounded
I have a computer science background. When I look at the optimizations an autoresearch loop produces, I can usually follow the technique. Cache-line alignment. Branch prediction hints. Loop unrolling. SIMD vectorization. Data structure substitutions. These aren't mysteries — they're patterns I studied.
But the cumulative effect is already harder to reason about.
After 200 iterations of an optimization loop, the code doesn't read like something a human wrote. It reads like something that was _grown_ — each change locally sensible, but the aggregate shaped by a fitness function rather than an architectural vision. The variable names are fine. The structure is legal. But the _why_ of each decision traces back to a measurement, not an intention.
This is manageable today because the techniques are recognizable. I can see what the agent did and understand the category of optimization, even if I wouldn't have thought to apply it there.
But I can see where this goes.
As models get better at optimization, the techniques get more exotic. Bit manipulation tricks. Memory layout optimizations that only make sense for a specific cache hierarchy. Instruction-level parallelism that a human would need hours to verify. Eventually — and I don't think "eventually" is very far away — the optimizations will be functionally opaque. Not wrong. Not buggy. Just beyond practical human review.
The review problem already exists for AI-generated code. Every team using agents is already merging code that no human fully traced. Autoresearch loops compound this: instead of reviewing one AI-generated implementation, you're reviewing the _result_ of hundreds of AI-generated modifications, each validated against a metric but none reviewed by a person.
The question isn't whether to accept this. It's how to build the trust infrastructure that makes acceptance rational:
- Tests that define the behavioral contract completely enough that you trust metric-validated changes.
- Benchmarks that are representative enough of real workloads that improvements transfer to production.
- Rollback mechanisms that make acceptance reversible.
- Monitoring that catches the cases where the metric improved but something else degraded.
Get comfortable with not fully understanding your own codebase. This is the direction. The alternative — insisting on human comprehension of every optimization — is insisting on leaving performance on the table. Permanently.
The New Primitive
CI/CD was a paradigm shift: automate testing and deployment. Run tests on every commit. Deploy on every merge. It took a decade to become standard, and it changed everything about how teams ship software.
Continuous autonomous improvement is the next layer.
Not just "test and deploy automatically" — _improve automatically_. Run optimization loops against your codebase the way you run linters: continuously, in the background, proposing changes that make things measurably better.
There's no reason this shouldn't always be running. The cost is compute. The benefit compounds. And the coordination model is already familiar — agents propose, tests validate, humans review at whatever cadence makes sense.
The stack becomes:
1. CI — verify that changes don't break things
2. CD — deploy verified changes automatically
3. CO — continuously optimize, validated by the same test infrastructure
The third layer is new. Most teams don't have it yet. But the early adopters I'm watching — and my own experience — suggest what it looks like when it clicks: software gets faster, smaller, and more efficient _without any human deciding to work on performance_. It just happens. As a property of the system, not a task on the roadmap.
At scale, autonomous optimization isn't one loop — it's a system of loops. An orchestrating agent
identifies optimization targets, spins up researchers against different metrics, validates that
improvements from one loop don't regress another, and synthesizes findings. This is agents
coordinating agents, and it works today.
The Paradigm
Here's what's actually changing.
For decades, optimization was a human activity. A skilled engineer profiled the code, identified bottlenecks, applied techniques from experience, and verified the improvement. It was artisanal — slow, expensive, and dependent on individual expertise.
Autoresearch makes optimization a _system design problem_. You don't optimize the code. You design the loop that optimizes the code. You choose the metric. You write the tests that define correctness. You set the constraints. Then you let it run.
The engineer's role shifts from _performing optimization_ to _designing optimization infrastructure_. From doing the work to defining what the work should achieve.
This is the same upstream migration that happened with overnight agents — from implementation to specification, from execution to judgment. But it goes further. With overnight agents, you're still reviewing individual changes. With autoresearch, you're reviewing _outcomes_. Did the metric improve? Do the tests pass? Does production look healthy?
The code itself? It's intermediate representation. It's the machine's concern.
That's uncomfortable. It should be. We've spent our careers treating code as the artifact — the thing we craft, review, discuss, and take pride in. The shift to treating code as a means rather than an end is not just a technical change. It's a professional identity change.
But the numbers don't care about our comfort. A 4.7x latency improvement doesn't care whether a human understands how it was achieved. It just serves requests faster.
The paradigm is this: define what good looks like, measure it continuously, and let the system move toward it.
That's not new as a principle. It's new as a practice — because for the first time, we have agents capable of executing the loop at scale, autonomously, overnight, across every measurable dimension of software quality.
The code gets better while you sleep. Whether you understand how is increasingly beside the point.