01 article

She shipped 2,000 PRs a month without reading the code

Lauren Tan ships thousands of pull requests a month with almost no human review because she built a verification loop that lets her agents check their own work. The approach is real and copyable, but it depends on a precondition most distributed systems cannot meet.

She shipped 2,000 PRs a month without reading the code

She shipped 2,000 PRs a month without reading the code

The New Stack reported last month that Lauren Tan, an engineer on the Grok team at xAI who shipped work at Cursor and Meta before, has around two thousand pull requests a month in production. Her own numbers are a little lower: a thousand in the last full month, pacing to double. Nobody has audited the PR list behind any of it, and I am flagging that up front because the number is the part you will remember. The part worth stealing is smaller. She almost never reads the code before it ships, and she says it herself in the talk, "I almost never look at code anymore." She can get away with it because she built the scaffolding that makes that safe.

The loop is infrastructure, not a prompt

Tan keeps a bundle of skills she calls pstack. The load-bearing part is a skill that, pointed at a repo, generates two things: a feature map of every area of the app, and a command-line interface that starts the whole application and hands back structured JSON. The JSON is the important bit. It is what lets an agent branch on a parsed field, start the app, inspect state, and decide for itself whether the change it just made behaves the way it was supposed to.

Her framing is blunt. To trust an agent's output, you either sit there and watch it, or you make it hand you evidence. The verification skill does the second, because it is a check the agent does not own and cannot quietly relax. She says she would unironically build her own debugging tools, or even pick a different tech stack, to get a loop that closes by itself. Tokens are cheap and getting cheaper every quarter. Your attention is the scarce thing, and it is what you want to spend at the highest leverage point, not on a four thousand line diff you will not understand anyway.

The catch: the app has to fit in one process

Here is the part the headline leaves out. The approach depends on an application that starts from a command line in seconds and can be thrown away once the agent is done with it. A frontend, a compiler, or a single service with a database all clear that bar. The moment your product is an order service plus a payments service plus a queue plus several databases, there is no such runtime by default, and supplying one that keeps up with hundreds of parallel agents becomes the actual engineering problem.

There are three ways teams currently fake it, and all three fail in a known way. Mocks are cheap and parallel, but a mock encodes what a dependency did the last time someone looked at it, and it drifts the instant the real service changes. The agent's loop closes cleanly, the failure is silent, and it shows up after merge. A full copy of the stack per change is faithful and isolated, but its cost scales with the number of services times the number of concurrent changes. Shared staging is faithful and cheap because there is only one of it, which means every agent is overwriting the others, and one broken change becomes everyone's failed test.

I would hold the mechanism at high confidence and the two thousand number at low. The mechanism you can test by building it. The number is one person's account, relayed once, with no second dataset to move it either way.

Lock the codebase down so the agent has fewer choices

The second pillar is the codebase itself. Tan's app is an Electron project she calls Dune, and the design goal is to cut down the number of decisions an agent makes on any task. She puts it as make the shortest path the correct one. A coding standard requires the agent to remember and obey every single time, but an architectural constraint removes the wrong option from what it can see. Dune does that with mechanical rules: feature state, components, and logic live in the same place; hard process boundaries keep heavy work off the UI; CI blocks cross-layer and cross-process imports; and lint or compile gates disable the patterns the agent keeps reaching for. Banning a particular hook and plain comments is their specific choice; the transferable move is to find the shortcut the agent misuses and seal it off with a boundary it cannot argue with.

She also keeps the feature map of the whole app, maintained in lockstep with the code, so an agent can figure out how something is supposed to behave before it starts guessing. That is the kind of documentation most of us avoid because it goes stale, but as navigation for a machine that re-learns the codebase from scratch every run, a stale map is cheaper than no map at all.

What I would and would not copy

I would steal the ordering. Build the verification infrastructure before you hand more of the work to the agent. In the loop, you approve every step. On the loop, you supervise a system that checks itself, and you only show up where the leverage is. That shift is the actual productivity, and it is available to a team of five, not just a solo engineer with a 24/7 cloud agent farm. I would not copy the throughput number into a slide, and I would not pretend the single-process precondition is a detail. If your product is a distributed system, the verification loop is still the right idea, but the runtime problem is yours to solve first, and nobody has published a clean answer for it yet.

The useful sentence from the whole thing is Tan's: an agent that can check its own output keeps working until the task is done. The rest is deciding how much you are willing to build so that a machine can hold itself accountable.

Comments