← All writing
Policy · · 9 min

Whoever gets there

On existential risk, the race, and a position I hold without enjoying it.

AI Culture

There are two arguments about superintelligence I keep running into, and the irritating thing is that I think they’re both correct.

The first says we’re building something nobody knows how to control. Not in a movie way. In a boring, technical, we-cannot-currently-specify-what-we-actually-want way. Alignment is an open research problem, and not the kind where everybody knows the shape of the answer and it’s a matter of grinding it out. Nobody at any of these labs will tell you with a straight face that they can reliably point a system substantially smarter than us at a goal and have it stay pointed. If that’s the situation, racing is the worst possible thing you could be doing.

The second says the United States is in a race with China whether it enjoys being in one or not, and whoever reaches the frontier first sets the terms for everyone who shows up later. Not just militarily. The defaults, the standards, the norms everybody else ends up building against. If that’s the situation, slowing down is the worst possible thing you could be doing.

And the usual move is to pick one and treat the other as bad faith. Safety people hear “but China” as a lobbying phrase, which, honestly, fair, it very often is. Race people hear “existential risk” as science fiction, or as regulatory capture wearing a philosophy degree, and that’s also fair, because sometimes it is. Both accusations land often enough to keep getting made.

I don’t think you get to duck it, though. So here’s where I’ve ended up: the US has to keep pushing, hard, and has to build the safeguards into the same motion rather than bolted on afterward, and when those two genuinely conflict (which they will, in specific budget lines and specific deployment calls, not in the abstract) I come down on staying ahead. I don’t feel great about typing that. Let me at least show my work.

The part where the doomers are right

The strong version of the slow-down argument isn’t “AI scary.” It’s about what competition does to margins. Safety margin is always the cheapest thing to cut and the last thing to visibly fail, which makes it the first thing to go when you’re behind and the quarter is ending. That’s not really a claim about AI, it’s a claim about organizations, and the base rates there are awful. Boeing. Deepwater Horizon. Every financial product that was fine right up until it wasn’t. I can’t name you an industry that held its margins voluntarily while a competitor was eating its lunch.

The other half is the one that actually gets me: this is the rare category of mistake with no iteration loop. I wrote a whole thing recently about how most work is worth doing badly, because the bad version teaches you things and you fix it later. Great for a side project. Completely useless here. “We’ll learn from the first one” isn’t available to you when the first one is also the last one, and I want to be upfront that this cuts against basically everything else I believe about building things.

So I’m not going to pretend the worried people are being silly. They’re not. A lot of them are the same people who built the thing, which I’d say is at least mild evidence they’ve thought it through.

But slow relative to what

Here’s where I get off, though. “Should we slow down” is never actually the question on the table. The question is “should we slow down relative to everyone else,” and those are wildly different questions with wildly different answers.

Refusing to build the thing is not the same as the thing not getting built.

Refusing to build the thing is not the same as the thing not getting built.

If the US throttles and nobody else does, the number of superintelligent systems in the world does not go down by one. It stays exactly the same. The only variable that moved is who’s holding it and how much say you have in what happens next. You haven’t prevented anything. You’ve recused yourself.

And the alignment work isn’t cleanly separable from the capability work in the way the tidy version of the argument needs it to be. Interpretability research needs frontier models to interpret. Evals need something worth evaluating. The safety techniques that have actually worked are, empirically, mostly downstream of somebody building the thing and then poking at it for a year. A country that opts out doesn’t become the world’s conscience on this, it becomes a commentator with strong opinions and zero leverage, and I don’t think anybody in history has been talked into a safety standard by a party who wasn’t in the room.

I want to be careful with the values pitch, because “American values” coming out of a country that does what this one does is not a clean sell and I’m not going to dress it up as one. I’m not claiming we’re the good guys. I’m claiming that on the narrow question of whether a researcher can publish a result that embarrasses the state, or whether a person can say out loud in public that the model got it wrong, the gap between here and there is not close. And those specific freedoms are really important for catching problems early. Safety mostly depends on somebody being able to say the uncomfortable thing without it ending their career. That’s not a flag-waving argument, it’s a pretty mundane one about error correction.

What I’d actually want

If the position is “go fast and build the safeguards into the same motion,” then the second half has to be real and not a paragraph in somebody’s press release. So, concretely, the things I’d spend money and political capital on:

  • Interpretability funded like a moonshot instead of like a compliance line item. The public money going into understanding what these systems are doing internally is a rounding error next to what’s going into making them bigger, and that ratio is more or less the whole ballgame.
  • Safety cases before deployment, with the burden sitting on the lab. Aviation and pharma both work this way: you demonstrate it’s safe, nobody has to prove it’s dangerous first. It’s slower, and it’s also why you board a plane without thinking about it.
  • Keeping the physical lead, because that’s the part policy can actually touch. Chips, sure, but increasingly energy and permitting. The frontier bottleneck is turning into gigawatts and transformers and interconnect queues, and none of that is a software problem.
  • Mandatory incident reporting and shared eval infrastructure, so four labs don’t each independently rediscover the same failure mode at 2 am and then quietly not mention it to each other.
  • A functioning visa system, which is the cheapest AI policy available to this country and the one we keep declining to have. Half the people you’d want building this are stuck in a lottery, on purpose, every year.

None of that is a brake. That’s the stuff that makes the lead worth having, because being first to a system nobody can trust isn’t winning, it’s just being first.

The part I can’t resolve

The honest shape of what I’m doing here is a bet, and the payoffs aren’t symmetric, and I know it. Wrong in my direction is potentially unrecoverable. Wrong in the other direction gets you a worse century, which is bad, but it’s a century. Anybody telling you those two are the same size is selling something.

Which means the whole thing rests on my probability for the catastrophic outcome being low enough that the expected value comes out where I’ve put it. And I want to say plainly that I can’t defend my number. I don’t really have one. I have a vibe with a decimal point stuck on the end of it. So does everybody else, including the people who publish theirs with error bars. If the real number is a lot higher than I think it is, then I’m wrong in the most expensive way a person can be wrong, and none of the geopolitics will matter, on account of there not being any.

What keeps me where I am is mostly that I don’t believe in the world where we stop. Not because stopping would be bad. Because I have never once seen a coordinated global slowdown on a technology with this much money and this much strategic value attached to it, and betting the future on one happening now, for the first time, reads to me less like caution and more like a different flavor of wishful thinking. Given that it’s getting built either way, I’d rather be in front, with the money and the leverage and the labs, than behind and hoping.

Anyway, I don’t like this position. It’s just the one I keep landing on, which isn’t the same thing as liking it, and I’d honestly love for someone to talk me out of it, though so far the attempts have mostly been people restating one of the two arguments I opened with 🫠. I have no security clearance and no lab access, and I am reasoning entirely from things published in public. Here’s hoping the whole thing turns out less dramatic than any of us think.

Read similar posts
9 min

Write it down once

Most of what makes Claude Code work for me isn't prompting, it's a few files I wrote once: a plan-mode habit, a /simplify pass welded onto /commit, a hook that won't let me hand-write a commit, and a CLAUDE.md that's mostly scar tissue.

7 min

Documentation for the you who forgot

I came back to a project after 14 months and found two notes I'd left myself, in the same file, in the same voice, and one of them saved me 40 minutes and the other cost me an hour and a half.