More

    OpenAI says its new ‘Astra’ AI can build attacks without any human help

    Published on:

    OpenAI says its new ‘Astra’ AI can build attacks without human help

    Astra is the first OpenAI model to reach its “Critical” cybersecurity threshold, meaning it can find previously unknown vulnerabilities and develop ways to exploit them across hardened systems.

    — OpenAI says its upcoming Astra model can autonomously discover previously unknown software flaws and turn them into working attacks, earning the company’s first “Critical” cyber capability rating.

    — In testing, Astra exploited known vulnerabilities, found two new flaws, escaped a hardened browser sandbox and combined operating-system weaknesses to gain root access.

    — OpenAI has delayed parts of Astra’s development to add safeguards and plans to limit its most advanced cybersecurity capabilities to selected testers, amid concerns that such tools could rapidly exploit crypto software flaws.

    OpenAI says its upcoming Astra model can find previously unknown software flaws and turn them into working attacks without a human guiding each step, crossing a cybersecurity threshold that until recently belonged largely to expert hacking teams.

    It is the first model OpenAI has classified as having “Critical” cyber capabilities under its Preparedness Framework, the firm wrote in a Tuesday post.

    To qualify, a model must be able to find previously unknown software flaws, known as zero-days, and develop working exploits for them across hardened real-world systems without human intervention, or devise and execute an attack from little more than a high-level goal.

    In testing, Astra scored 100% on a benchmark for developing exploits from known vulnerabilities and found two previously unknown flaws while building an exploit chain on a separate internal test.

    It also broke out of a hardened browser sandbox and executed commands on the host computer, while separately finding and combining multiple flaws in an operating system to gain root access, OpenAI said.

    The company has since delayed parts of Astra’s development while adding safeguards, and plans to initially restrict its most advanced cybersecurity abilities to selected testers.

    That capability is particularly relevant to crypto, where a software flaw can be converted into money within minutes. CoinDesk reported in June that increasingly capable AI models could compress the work of searching code, finding misconfigurations and assembling attacks from days or weeks into machine-speed operations.

    At the time, security researchers said the bigger change was not necessarily a new class of hack, but how quickly existing weaknesses could be found and exploited.

    The advance follows other signs that frontier models are moving beyond answering questions and writing code. Anthropic’s Claude Fable 5 helped solve an 87-year-old mathematics problem in July, as CoinDesk reported.

    Anvil is a shared on-chain collateral layer built on a programmable letter of credit: reserve assets as a guarantee -no loan, no interest, keep custody & yield.

    Why it matters:

    Anvil is a shared on-chain collateral layer built on a programmable letter of credit: reserve assets as a guarantee -no loan, no interest, keep custody & yield.

    Related