Member-only story
The True Safety Approach of Modern AI Engineering
Boring Things Save Lives
Back in 2023, hundreds of the most influential people in the field of artificial intelligence signed a one-sentence statement: “Mitigating the risk of extinction from AI should be a global priority alongside other societal-scale risks such as pandemics and nuclear war.” The signatures include AI heavyweights such as Geoffrey Hinton, Yoshua Bengio, Sam Altman, Demis Hassabis, Dario Amodei. The CEOs of OpenAI, Google DeepMind, and Anthropic put their names next to the word extinction.
So if the people running the labs are worried enough to sign that, a fair question is: what are they actually doing about it? Not the safety blog posts. The real, day-to-day mechanism that’s supposed to keep a frontier model from going off the rails.
In the documentary The AI Doc: Or How I Became an Apocaloptimist, Sam Altman explained the very complicated process that the models go through to ensure safety much better than I could. Are you ready for this? Here are the steps:
- You build the model.
- You test it.
- You release it to a small number of people, and watch what happens.
- If it looks okay, you let more people in.
