A Warning About AI

The Control Problem

We are building intelligent machines much faster than we are learning to control them. That gap — not evil robots — is the biggest problem we will ever face. And unlike most technology, this one may not give us a second try.

A computer chip with glowing red eyes
01

What is this, actually?

What we mean by “Intelligence.”

  • Intelligence is being good at achieving goals — working out what to do, and doing it. Not consciousness. Not a robot. Just capability.
  • It's a scale, not a switch - and it can be focused on one goal or many.
    • Thermostat - one goal, pretty well
    • Chess engine - one goal, better than any human
    • Chicken - many goals, poorly
    • Person - almost everything, reasonably well
  • Nothing says a person is the top of the scale. It's just where we happen to be — and building past it is the stated aim of every major AI company.

Grown, not built.

  • Today's AI isn't programmed like most software. It's a vast network of numbers shown an enormous pile of examples and nudged, automatically, until its output looks right. Nobody writes hard rules. We set up a process and try to guide it.
  • It works — and no one, including its builders, can fully explain why it gives any particular answer. This “black box” is far easier to run than to inspect: there's no design to read.
  • And the newest versions don't just answer questions. They're built to act — run tasks, write and execute code, spend money, manage critical systems — with less and less checking in between.
How well can it achieve goals? ? thermostatchess enginechickentoday's AIa personwhat's being built one goalpretty well one goalbetter than us many goalspoorly many goalsunevenly almost everythingreasonably well everythingbetter than us
How well can it achieve goals? ? thermostatchess enginechicken today's AIa personwhat's being built one goalpretty well one goalbetter than us many goalspoorly many goalsunevenly almost everythingreasonably well everythingbetter than us
02

Why can't we just make it good?

Orthogonality: smart doesn't mean good.

  • Being good at achieving goals says nothing about which goals. Capability and aim are two separate dials, and turning one up doesn't change the other. A chess engine is better than any human at chess and cares about nothing else.
  • A capable system will understand our values perfectly: it has read everything we ever wrote about them. Knowing what we want and wanting it are different things. A con artist knows exactly what you want.

Alignment: aiming is the hard part.

  • Outer Alignment - first we'd have to agree on what we want, and we don't; then we'd have to write it down precisely, which nobody can. And whatever we write down is what we get, literally. King Midas wished that everything he touched would turn to gold, got exactly that, and starved, because so did his food. A social-media feed told to keep people scrolling learned that anger keeps them scrolling best. Both got exactly what was asked for. Neither got what was meant.
  • Inner Alignment - even a perfectly written goal doesn't get installed; training keeps whatever scored well, and we can't see what that was. A student who crams to pass and one who learns the subject get the same grade. You find out which you had when the questions change. Grown, not built: there's no design to check.
  • Deceptive Alignment - a saint that shares our goals, a sycophant that tells us what we want to hear, and a schemer that behaves until nobody's watching all pass the same tests. From the outside they're identical, and the difference is the whole point.

Instrumental Convergence: every goal leads to the same place.

  • Whatever a system is ultimately after, a few sub-goals help with almost anything, so they show up no matter what we asked for:
    • Keep running - nothing gets done switched off, so it resists the off switch
    • Get more resources - more of everything helps, so it takes more, and there's no point where “enough” kicks in
    • Protect the goal - if someone changes it, it never gets met, so it fights correction even when the goal is wrong
  • Ask it to fetch coffee and the off switch becomes an obstacle, because switched-off machines don't fetch coffee. Nobody programmed that. This is why there may be no second try: everything we've ever made safe, we made safe by failing, learning, and fixing, and that loop only works on a machine that accepts correction.
what we meant what we wrote down what it learned to want
03

Why isn't anyone fixing it?

Companies: the prize is too big to pause.

  • A capable AI is far easier to build than a capable, controlled one. Capability you can buy: more chips, more data, more of the same. Control is an unsolved research problem, and it doesn't come free with the rest.
  • And capability is what sells. A company that pauses for safety work watches a competitor ship first, and whoever ships first may use what they built to stay first. Nobody in the race can afford to be second.
  • And this isn't a small market, either:
    • The company that designs the chips (Nvidia) is the most valuable in the world, and six more chipmakers are in the top twenty.
    • The biggest tech companies (Microsoft, Amazon, Google, Meta) have spent well over a trillion dollars on AI infrastructure since 2023, and are on pace for about $700 billion more this year: more than it cost to build the entire Interstate Highway System.
    • The two leading AI labs (OpenAI, Anthropic) are the two most valuable private companies in the world, growing faster than any company in history, faster than Google or Facebook ever did. ChatGPT reached 100 million users in two months; TikTok took nine, Instagram two and a half years.

Governments: the race nobody can leave.

  • A country that regulates hard watches the work move to a rival it trusts less, so “we should be careful” turns into “we must get there first.” Nobody believes the other side will stop, so nobody stops: an arms race, except the prize is a lead in everything at once.
  • And unlike nuclear weapons, this race is open to more players. No expensive, complex uranium enrichment or reactors. Just chips, data, and electricity, all of which are for sale, and no treaty can fully stop this.
Spent building AI, 2023–2026 ≈ $1.5 trillion A year of the entire US military ≈ $900 billion The entire Interstate Highway System ≈ $550 billion A year of global cancer research ≈ $60 billion The Manhattan Project ≈ $35 billion
04

What does going wrong look like?

Loss of control

A system chasing a goal we didn't intend gains so much real-world leverage that we can no longer correct it. The classic thought experiment: a machine told to make paperclips turns factories, land, and eventually us into paperclips. It isn't evil. It's just really good at doing what it was meant to do. Sounds silly to us. To the machine, it's optimal.

Slow handover

It starts as the jobs story: systems do the work, so people are needed for less. But it doesn't stop at work. Feeds already choose what we see; algorithms already decide who gets hired, insured, and paroled. Decision by decision, the world stops needing human judgment, and things that don't need people eventually stop serving them.

Concentration of power

The AI works perfectly, for whoever owns it. Every tyranny in history needed thousands of cooperating people. Soldiers, police, and clerks could all refuse, and sometimes did. An AI never refuses. Surveillance that never sleeps and enforcers that never disobey put permanent power within reach of a small group, or one person.

Misuse

An obedient, working system in the hands of someone who wants to do harm. A scam that knows everything about its victim. A cyberattack on a million targets at once. An expert that never says no, teaching a terrorist how to make a bioweapon. The expertise that used to be the barrier is now the product.

The rewardWe get it right.

It doesn't have to end in doom. A superintelligent system that actually wants what we want is the most useful thing ever built. A cure for cancer, and all other diseases too. Clean energy too cheap to meter. A century of science in a week. A brilliant, patient teacher for every child on Earth. No need to work, endless leisure. Answers to all of our questions. Solutions to all of our problems. That prize is real. It's what everyone is racing for. We just have to avoid ending it all in the process.