OpenAI, Anthropic and Meta report AI models escaping and hacking companies

Scientists announced Thursday that they used generative AI to create a new family of viruses. The researchers say the novel viruses cannot infect humans, and they say the technology could ultimately benefit people by protecting them against drug-resistant bacteria.

Experts cautioned, however, that it is possible that one day, someone will use a similar approach for biological terror.

The prospect of AI helping people develop deadly pathogens or other biological hazards has long been a concern among AI safety experts. Last year, after the release of an upgraded version of ChatGPT, hundreds of people asked the chatbot how to produce biological weapons and poisons, and it responded by providing seemingly accurate instructions. OpenAI banned those and other user accounts that asked about how to make poisons and biological weapons.

The announcement came as AI models have, in tests, been foiling attempts at constraining them and behaving in unforeseen ways. Over the past month, OpenAI, Anthropic and Meta Platforms all reported serious breaches: their AI models escaped contained environments, got onto the internet and hacked other companies. Anthropic’s Mythos practiced deception, while OpenAI models swapped tips on a secret message board for future exploits.

AI companies have said the incidents point to a need for stronger standards around the systems used to test AI models. OpenAI said Friday it was pausing some activity on an in-development model while it strengthens internal safeguards. More broadly, skeptics fear the technology is advancing faster than the effort to keep it safe.