AI Builders
Lauren Tan on Trusting Coding Agents: Verification, Feature Maps and Constraints That Bite
Lauren Tan, a React Compiler author now at Cursor, went from watching every agent output to agents merging their own pull requests. The path: verification that closes the loop, feature maps, skills tested with evals, and constraints that fail the build.
Akmal Alif · 8 October 2026 MYT

AI Builders is where we follow engineers and tech creators who explain, in detail, how they actually use AI in their work. The first is Lauren Tan, a former member of Meta's React core team and one of the authors of React Compiler 1.0 (Tan et al., 2025), who joined Cursor in 2026. Cursor is now part of SpaceX, which agreed to buy the company in June 2026 and folded it in alongside its AI unit, SpaceXAI (O'Kane, 2026).
The recording below is a live session from around August 2026 in which Lauren takes questions from a host and a chat audience. It was uploaded by a third-party channel whose title and description add claims that are not in the talk, so this post follows only what Lauren says. The timestamps jump to the matching moment.
The problem is trust
Lauren starts with a caution. The best use of AI, in Lauren's view, is as a pair programmer, not as a way to replace writing code or to outsource your thinking. Agents save time on writing but can cost more in review, and the pressure to ship a hundred pull requests a week without understanding them is real (watch from 0:00).
The theme of the session is trust. Agents guess, hallucinate and announce that they have found the cause of a bug for the hundredth time, and every false alarm costs some of an engineer's trust (watch from 1:02). Lauren compares it to management: if you do not trust your team, you end up micromanaging them, and you cannot delegate more work than you can watch (watch from 2:03).
Lauren draws this as a curve. A year ago it meant being deeply in the loop with a handful of agents, watching every output. Now Lauren's agents merge their own pull requests: one morning, about twenty had landed overnight, and Lauren reviewed them on main (watch from 3:05). The productivity chart that follows the same curve shows roughly a thousand merged PRs in the previous month (watch from 4:06; 5:06). Lauren is the first to say that the volume invites the question of how good that code is, which is what the rest of the session answers.
Verification closes the loop
The single most important skill, Lauren says, is verification: giving the agent a way to run the code for real, whether that means taking CPU traces, capturing heap snapshots or opening an iOS simulator, in whatever way users reach the application. Verification does not guarantee good code, but it lets the agent at least produce correct code, and that is the step that makes trust possible (watch from 6:08).
The story behind it comes from Lauren's first week at Cursor, doing performance work on the new agents window with Chrome DevTools traces and an agent that understood the code base no better than Lauren did. Without verification, you are the verifier: you open the build, copy screenshots and console errors back to the agent, and become the bottleneck that stops any parallel work (watch from 8:15). Lauren's first skill taught agents to drive the app through the Chrome DevTools Protocol, the same idea that works for Electron, web or iOS apps (watch from 9:15).
Feature maps: teach the agent the product
Being able to run the app was not enough; the agent did not know what anything in it was. Told that “the left sidebar is laggy”, it would wander through the code trying to find the sidebar (watch from 10:15). The fix was a feature map: a file that tells the agent how to reach every feature from the user's point of view, including keyboard shortcuts and the DOM attributes to select (watch from 11:16).
That turns even a terrible bug report into something an agent can act on. Lauren's example is a feedback channel where a report can be a screenshot followed by three question marks (watch from 12:16). These skills now live in Pstack, Lauren's plugin of skills, which includes one skill to create a verification setup and another to keep it up to date (watch from 13:16).
Skills from failure modes, evals for skills
Pstack was never planned. Lauren built it skill by skill from observed failures, such as an agent that confidently named a cause without having read the code that mattered (watch from 14:19). The management analogy returns: a skill is how you onboard a strong engineer who arrived five seconds ago with no context, and it is just Markdown (watch from 15:20).
To keep skills honest, Lauren treats evals as unit tests for agents (watch from 17:21). A coordinator agent writes a rubric and spawns sub-agents in neutrally named directories, because agents that can tell they are being evaluated change their behaviour. The skill is run across several models, a judge from a different model checks the scoring, and the loop repeats until the score stops improving (watch from 18:22; 21:25). The human job, Lauren says, is to be a very good backseat driver: read the tool calls and the reasoning, find where agents fail, and turn that into the next skill (watch from 20:24).
From local to cloud, one rung at a time
Lauren's image for the new role is a head chef: no longer cooking every dish, but designing the kitchen and assigning the stations (watch from 22:25). The practical advice is to start locally, where you can watch an agent use your app, and only then move to cloud agents (watch from 23:25).
Once verification works, it pays off for the whole company. Cursor runs an agent that takes incoming bug reports, launches the app on its own cloud machine with the same control skills, and tries to reproduce each one; in Lauren's example it reported that a bug was real but already fixed on main (watch from 25:25). The warning is not to skip rungs. Jumping straight to thousands of cloud agents before you trust one mostly burns tokens (watch from 26:28).
Greenfield code needs the strongest guardrails
The most surprising argument is about where agents are dangerous. Large company code bases already carry guardrails built for thousands of engineers of varying skill, and agents benefit from them (watch from 30:31). The bigger risk is a fresh, “vibe-coded” prototype with no constraints at all, where agents take whatever shortcut solves today's task and the code base spirals beyond anyone's understanding (watch from 31:31).
That was the state of Grok Bot, the agent app SpaceXAI launched in August 2026 (Clover, 2026), whose prototype had been built fast with nobody reading the code (watch from 32:32). Lauren rebuilt it on a new, deliberately strict architecture over more than 600 pull requests, and only after that investment stopped reading every line (watch from 33:34).
Constraints that bite
The internal architecture, which Lauren describes as something like Next.js for Electron apps designed for agents to write, is enforced by CI with checks that would irritate a human (watch from 37:41):
React's useEffect hook is banned. It is one of React's most common sources of bugs.
Code comments are banned. Agents mostly write comments about the history of a change, such as a reviewer's one-off remark, which then reads like a permanent rule (watch from 38:42).
Process boundaries are checked. Electron code for the main process and the renderer lives in separate directories, and CI inspects the dependency graph so heavy work cannot slip onto the thread that has 16 milliseconds to draw each frame (watch from 40:43).
Every feature lives in one directory. Agents copy existing patterns, so the conventional path is made the shortest path (watch from 42:44).
Lauren ranks enforcement in layers (watch from 44:46). Code conventions, lint rules, compiler diagnostics and CI checks are hard: they turn the build red. Rules files, skills, review bots and style guides are soft: agents can forget them. Relying only on the soft layers, Lauren warns, means the code base degrades sooner or later. That is also why a strict compiler, such as Rust's, builds confidence (watch from 45:48). The worst place to be is enforcing your invariants by hand in code review; every repeated review comment should become a lint rule, a CI failure or a design that makes the mistake impossible (watch from 46:48).
The cost question
Asked whether this works on an ordinary token budget, Lauren is candid: working at an AI lab means effectively unlimited tokens, so not everyone should copy the approach exactly (watch from 47:49). The framing offered for engineering leaders is return on investment. Refactoring a code base until even modest agents write good code in it costs a lot of tokens up front, but it can be cheaper than the people and time it replaces, and it keeps a team small (watch from 48:51; 49:52).
The payoff Lauren describes is that designers, product managers and other non-engineers can now ship changes safely, because the constraints catch what they would miss (watch from 53:57).
What to take away
Trust is earned with verification. Give agents a way to run and observe the software before giving them more autonomy.
Teach the product, not just the code. A feature map lets an agent connect a vague report to the right screen.
Turn every failure into a skill, and test the skills. Evals keep instructions honest across models.
Prefer hard constraints to soft ones. If you keep writing the same review comment, encode it as a rule that fails the build.
Climb the curve in order. Local, then cloud, then automation; skipping rungs wastes tokens and trust.
References
Clover, J. (2026, August 11). Grok Bot brings always-on AI agents to macOS and iOS. MacRumors. https://www.macrumors.com/2026/08/11/grok-bot-macos-ios
O'Kane, S. (2026, June 16). SpaceX to acquire Cursor for $60B in stock, days after blockbuster IPO. TechCrunch. https://techcrunch.com/2026/06/16/spacex-to-acquire-cursor-for-60b-in-stock-days-after-blockbuster-ipo/
Tan, L., Savona, J., & Zhang, M. (2025, October 7). React Compiler v1.0. React Blog. https://react.dev/blog/2025/10/07/react-compiler-1
Vision Sabbat. (2026, September 14). Lauren Tan – SpaceXAI engineer [Video; third-party upload of a live session]. YouTube. https://www.youtube.com/watch?v=EWSUvEyFwjc