OpenAI said its upcoming Astra model can autonomously discover previously unknown software vulnerabilities and turn them into usable attack methods, making it the company’s first model to reach the Critical cybersecurity capability threshold. According to ChainCatcher, OpenAI said that under its Preparedness Framework, this level means a model can find zero-day flaws in hardened real systems and develop working exploit code without human intervention, or design and carry out an attack from a high-level objective alone.
In testing, Astra scored 100% on a benchmark for developing exploit code from known vulnerabilities and found two previously unknown flaws in another internal test. The model also broke out of a hardened browser sandbox, executed commands on the host machine, and gained root access by combining multiple operating system weaknesses.
OpenAI said it has delayed part of Astra’s development to add more safety protections and plans to restrict access to its most advanced cybersecurity capabilities to screened testers only. The company said the capability is especially relevant to the crypto industry, where software vulnerabilities can be turned into financial losses within minutes.