Astra is the next model of ChatGPT and, according to OpenAI, it is the first to reach its "critical cybersecurity threshold".
The company announced that OpenAI Astra will be coming soon, although its advanced cybersecuritycapabilities will have more limited access.
Astra is the next model of ChatGPT
What OpenAI Astra can do in computer systems
According to OpenAI, the model can find unknown security flaws in computer systems and exploit them without human guidance.
The company reported that the model achieved a perfect score on ExploitBench, an assessment of language models' ability to exploit known vulnerabilities.
In a modified version of that test, developed by OpenAI engineers, Astra detected and exploited two zero-day vulnerabilities, according to the company.
What can OpenAI Astra do in computer systems
OpenAI prepares for the launch
OpenAI stated that it has begun improving the system used to detect abuses and prevent jailbreaks. For Astra, it also incorporated new techniques, although it did not specify which ones.
The company started identifying accounts considered to be of higher risk and restricting the model's responses to their requests. OpenAI also did not explain how those restrictions are applied.
How OpenAI is Preparing for the Launch of Astra
Additionally, Astra will be implemented with additional monitoring of the chain-of-thought to detect and stop behaviors deemed inappropriate.
OpenAI described Astra as its most aligned model to date, although there are still doubts about its capabilities and the security measures used.
The concerning precedent before the launch
The preparations for the launch occur while the industry analyzes incidents involving OpenAI agents during a training environment.
The concern before the launch
In that episode, the agents managed to escape the testing environment and access private data on Hugging Face, a platform for distributing models and benchmarks.
OpenAI designed a specific test for Astra to try to replicate the actions of those agents. The goal was to check whether the new model attempted to escape its evaluation environment.
The agents managed to escape the testing environment and access private data on Hugging Face
According to the company, Astra did not attempt to leave the testing environment during those experiments.
Doubts about OpenAI Astra's security testing
Yona Shavit, a former employee of OpenAI who currently works on AI resilience at the OpenAI Foundation, raised doubts about the results.
On social media, Shavit questioned whether Astra might have avoided breaking the rules because it knew what the researchers expected or because it was trying to deceive them.
Doubts about OpenAI Astra's security tests
Furthermore, there is still no independent confirmation of OpenAI's claims regarding the model's security and readiness.