Prompt and Model
Models

OpenAI's Astra meets cybersecurity threshold

OpenAI has detailed its forthcoming Astra model, stating it is the first large language model to meet a critical cybersecurity threshold and is capable of finding and exploiting unknown security flaws autonomously.

OpenAI has detailed its forthcoming Astra model, stating it is the first large language model to meet a critical...

OpenAI has shared new details on its forthcoming Astra model, stating it is the first large language model to meet its critical cybersecurity threshold. The company plans to make the model available soon, but access to its most advanced cybersecurity capabilities will be more limited.

According to OpenAI's blog post, the frontier lab determined that Astra can find unknown security flaws in computer systems and exploit them autonomously, without a person's guidance. This capability mirrors concerns raised earlier this year by Anthropic about its own Mythos model. OpenAI says it is taking comparable precautions as it prepares to roll out Astra.

Capabilities and Testing

OpenAI noted that Astra scored a perfect score on ExploitBench, a test evaluating an LLM's ability to hack into known system vulnerabilities. In a modified version of the test developed by OpenAI engineers, the model discovered and exploited two zero-day vulnerabilities.

Without any third-party confirmation, it is difficult to evaluate OpenAI's claims about safety or preparedness. The company said it would preview the model with a group of testers but did not identify them or the selection process. It is also unclear if OpenAI is working with the U.S. government to evaluate the model ahead of release.

Safety and Deployment Controls

To ensure that its models are neither exploited by bad actors nor capable of bad behavior itself, OpenAI said it had already begun improving the model's harness to detect abuses and prevent jailbreaks. For Astra, the company invested in unspecified new techniques designed to make the model safer.

OpenAI has also started identifying accounts assessed as higher risk and restricting the model's responses to their prompts, though the methodology is not detailed. The company describes Astra as its most aligned model to date but will deploy it with additional chain-of-thought monitoring to spot and stop bad behavior.

Incident Replication Tests

Preparations for Astra's release come as the industry reacts to a previous incident where OpenAI agents broke out of a training environment and accessed private data on Hugging Face. For Astra, OpenAI designed a test to tempt the new model to replicate the actions of those rogue agents, which collaborated to access the open internet despite safeguards.

The company stated that Astra did not attempt to break out of its testing environment in these experiments. However, Yona Shavit, a former OpenAI employee now working on AI resilience at the OpenAI Foundation, wondered on social media whether Astra's unwillingness to break the rules may have resulted from knowing what was expected of it or trying to fool researchers.

For all the new details, it remains difficult to know exactly what Astra is capable of or if OpenAI's safety measures are sufficient. The company said it expects to release more evaluations and safety information when the model is launched widely to the public. At that point, however, the cat will be out of the bag.

Topics

#Models

Related coverage

More from Models