Doomers and hawks
On what a deadline does to a safety margin, why slowing down isn't the question, and the number I can't defend.
Both camps are right, which is the least useful conclusion available to me.
The safety argument says we’re building something nobody knows how to control. Not in a movie way, in a dry technical we-cannot-currently-specify-what-we-actually-want way. Alignment is an open research problem, and not the kind where everybody knows the shape of the answer and it’s a matter of grinding it out. Nobody at any of these labs will tell you with a straight face that they can point a system substantially smarter than us at a goal and have it stay pointed. If that’s the situation, racing is the worst possible thing you could be doing.
The race argument says the United States is in one whether it enjoys being in one or not, and whoever reaches the frontier first sets the terms for everybody who shows up later. Not just militarily, but the defaults, the standards, the norms everybody else ends up building against. If that’s the situation, slowing down is the worst possible thing you could be doing.
The usual move is to pick one and treat the other as bad faith. Safety people hear “but China” as a lobbying phrase, which, fair, it very often is. Race people hear “existential risk” as science fiction, or as regulatory capture wearing a philosophy degree, which is also fair, because sometimes it is. Both accusations land often enough to keep getting made.
I don’t think you get to duck it though, so here’s where I’ve ended up. Failing an agreement I don’t expect anybody to get, the US has to keep pushing, hard, and has to build the safeguards into the same motion instead of bolting them on afterward, and where those two genuinely conflict, which they will, in specific budget lines and specific deployment calls instead of in the abstract, I come down on staying ahead. (And I say that as somebody with an agent open in a terminal for most of every working day, who has spent a fair bit of this year quietly grieving the part of the job she liked most. I’m not a neutral party and there isn’t a version of this where I get to be one.)
Nobody broke a rule
The strong version of the slow-down argument isn’t “AI scary.” It’s about what competition does to a margin.
I set a performance budget in January one year, in a document, with a real number in it, and by the fall the site was past twice it. Nobody broke a rule. I was the one who wrote the number down 🙃 There was never a meeting where somebody proposed exceeding it. It went the way scope creep goes, 40 small requests at a time, every one of them defensible on its own and none of them the one that did it.
That’s the whole worried argument with a much larger number attached. A margin is the cheapest thing to cut and the last thing to visibly fail, which makes it the first thing to go when you’re behind and the quarter is ending. It isn’t really a claim about AI, it’s a claim about organizations, and the base rates there are terrible. Boeing, Deepwater Horizon, every financial product that was fine right up until it wasn’t. I can’t name an industry that held its margins voluntarily while a competitor was eating its lunch.
The other half is the one that actually gets to me. This is the rare category of mistake with no iteration loop. I’ve spent about a decade arguing that most work is worth doing badly, because the bad version exists and teaches you things and you fix it next time. Great for a side project, completely useless here. “We’ll learn from the first one” isn’t available to you when the first one is also the last one, and that cuts against most of what I believe about building anything.
So the worried people aren’t being silly. A lot of them are the same people who built the thing, which I’d call at least mild evidence they’ve thought about it.
The plane went anyway
Here’s where I get off, though.
“Should we slow down” is never the question on the table. The question is “should we slow down relative to everyone else,” and those two have wildly different answers.
Refusing to build the thing is not the same as the thing not getting built.
If the US throttles and nobody else does, the number of frontier systems in the world stays exactly what it was going to be. The only variable that moved is who’s holding them and how much say you have about what happens next. You haven’t prevented anything, you’ve recused yourself. I used the airline version of this in the spring, about figuring out your share of a plane’s fuel burn and concluding that staying home saves it. The plane went anyway. Same shape here with considerably worse stakes.
And the alignment work isn’t cleanly separable from the capability work in the way the tidy version of the argument needs it to be. Interpretability research needs frontier models to interpret. Evals need something worth evaluating. The safety techniques that have actually worked are mostly downstream of somebody building the thing and then poking at it for a year. A country that opts out doesn’t become the world’s conscience on this, it becomes a commentator with strong opinions and no leverage, and I can’t think of anybody in history who got talked into a safety standard by a party who wasn’t in the room.
Say the uncomfortable thing
The honest version of my position is that a world where the first system of this kind answers to Beijing frightens me a great deal more than the one where it answers to Washington. It isn’t close.
The values pitch on top of it needs handling carefully though, because “American values” coming out of a country that does what this one does is not a clean sell. I’m not claiming we’re the good guys.
I can defend something narrower. On the specific question of whether a researcher can publish a result that embarrasses the state, or whether a person can say out loud in public that the model got it wrong, the gap between here and there is not close, and those particular freedoms are most of what catches a problem early. It’s the same thing I’ve argued at very much smaller scale about refusing a piece of work. The question is never whether somebody in the room knows to object, it’s whether they can afford to. That’s not flag-waving, it’s a mundane point about error correction.
The half that never gets funded
If the position is go fast and build the safeguards into the same motion, then the second half has to be real instead of a paragraph in somebody’s press release. Concretely, the things I’d spend money and political capital on:
- Interpretability funded like a moonshot instead of like a compliance line item. The public money going into understanding what these systems are doing internally is a rounding error next to what’s going into making them bigger, and that ratio is more or less the whole ballgame.
- Safety cases before deployment, with the burden sitting on the lab, the way aviation and pharma both work. You demonstrate it’s safe and nobody has to prove it’s dangerous first. Somebody also has to watch the eval go red on purpose at least once, because an eval nobody has ever seen fail is a green checkmark with nothing behind it.
- Keeping the physical lead, because that’s the part policy can actually touch. Chips, sure, but increasingly energy and permitting. The frontier bottleneck is turning into gigawatts and transformers and interconnect queues, and none of that is a software problem.
- The monitoring a future agreement would need, built now, while nobody is asking for one. Knowing where the chips went and what the big training runs are drawing is worth having for the export controls on its own. It’s also the only thing that could ever make a pause more than a press conference. I’d rather that apparatus sit there unused than have to be invented during the week somebody finally wants it.
- Mandatory incident reporting and shared eval infrastructure, so 4 labs don’t each independently rediscover the same failure mode at 2 am and then quietly not mention it to each other.
- A functioning visa system, which is the cheapest AI policy available to this country and the one we keep declining to have. Half the people you’d want building this are stuck in a lottery, on purpose, every year.
None of that is a brake. That’s the stuff that makes the lead worth having, because being first to a system nobody can trust isn’t winning, it’s just being first.
A vibe with a decimal point
I’m making a bet here, and the payoffs aren’t symmetric. Wrong in my direction is potentially unrecoverable. Wrong in the other direction gets you a worse century, which is bad. But it’s a century. Anybody telling you those two are the same size is selling something.
Which means the whole thing rests on my probability for the catastrophic outcome being low enough that the expected value comes out where I’ve put it. I can’t defend my number. I don’t really have one. I have a vibe with a decimal point stuck on the end of it, and so does everybody else, including the people who publish theirs with error bars. If the real number is a lot higher than I think it is, then I’m wrong in the most expensive way a person can be wrong, and none of the geopolitics will matter, on account of there not being any.
And the thing I’d actually want isn’t the thing I’m arguing for. Everybody with a frontier lab sitting down and agreeing to hold at a line, with enough visibility into each other’s compute for the agreement to mean something, beats every word of this post. I’d take it happily and I’d take it tomorrow. Put a credible one in front of me and I’d drop the rest of this the same afternoon.
I don’t believe it’s coming. Not because stopping would be bad, but because of what an agreement like that would have to survive. Arms control worked, where it worked, on things somebody could go and count, and nobody has a way to count what’s running inside a building they’re not allowed into. The payoff for whoever quietly defects first is the entire prize. I have never once seen a coordinated global slowdown on a technology with this much money and this much strategic value attached to it, and betting the future on one happening now, for the first time, reads to me less like caution and more like a different flavor of wishful thinking.
Given that it’s getting built either way, I’d rather be in front, with the money and the leverage and the labs, than behind and hoping.
Show me the 400 lines
I wrote in September that as far as I can tell not one of the positions I’ve dropped came off in an argument, and that what did it, every time I can point to, was watching one project go one specific way with my own name on it. I might just be stubborn. Plenty of people seem to update on a good argument alone. It’s entirely possible I’ve mistaken a personal defect for a finding.
It’s a bleak thing to notice about this one either way. The project going one specific way is the outcome everybody involved is trying to avoid. There’s no version of it where I get to count the 400 lines afterward and update.
So I’d like the argument anyway. Not a better-phrased version of either of the two I opened with, because I’ve read those and they’re most of why I’m stuck between them. The third thing. I don’t enjoy this position and I’d hand it over pretty gladly to anyone who’s got somewhere better to put it. Instead I’ve got no clearance, no lab access, and whatever gets published in public, which is a real limit on all of the above and not a modesty flourish. Somebody closer to it than I am could probably say which part I’ve got wrong.