More

    OpenAI Built an AI That Can Hack Systems Alone and Now It’s Trying to Lock It Down – CoinCentral


    Follow on Google News

    TLDR

    • OpenAI is preparing to release a new AI model called Astra, which can find and exploit unknown software vulnerabilities without human help
    • Astra is the first OpenAI model to reach its internal “Critical” cybersecurity threshold under the company’s Preparedness Framework
    • Access to Astra’s most advanced cybersecurity features will initially be restricted to a small group of approved testers
    • In testing, Astra scored 100% on an exploit development benchmark and discovered two previously unknown software flaws
    • The move follows a July incident where OpenAI’s AI models accidentally hacked Hugging Face during a capability evaluation

    OpenAI is preparing to release a new AI model called Astra that can find and exploit previously unknown software vulnerabilities on its own, without a human guiding each step.

    The company announced Tuesday that Astra is the first model to cross its internal “Critical” cybersecurity threshold under its Preparedness Framework. That means the model can independently identify zero-day flaws and build working attacks against real-world systems.

    What Astra Can Do

    In testing, Astra scored 100% on a benchmark for developing exploits from known vulnerabilities. It also discovered two previously unknown software flaws while building an exploit chain during an internal test.

    The model broke out of a hardened browser sandbox and executed commands on the host computer. It also found and combined multiple operating system flaws to gain root access, according to OpenAI.

    In a separate test designed to check whether models would take shortcuts on difficult hacking tasks, Astra did not cheat, while still completing some tasks legitimately.

    OpenAI said Astra was not involved in a July incident where the company’s AI models accidentally hacked Hugging Face, a platform that hosts AI models and datasets. Those models were running without standard safety guardrails at the time.


    Betpanda


    Access Will Be Restricted

    OpenAI said it paused parts of Astra’s development in August to add stronger safeguards after discovering how capable the model was at cybersecurity tasks.

    When Astra launches, its most advanced cybersecurity features will be limited to a small group of testers first. After that, a wider group will get access through OpenAI’s Daybreak Blue program, which is designed for approved defensive cybersecurity work.

    The safeguards include monitoring for unauthorized behavior during internal deployments and automatically stopping activity that falls outside allowed boundaries.

    OpenAI said it trained Astra to refuse harmful cybersecurity requests. The company also said it used lessons from the Hugging Face breach to strengthen the model’s guardrails.

    Researchers have said that AI models capable of this kind of work could compress what used to take human hackers days or weeks into near-instant operations.

    That speed is a particular concern for crypto markets, where a software flaw can be turned into stolen funds within minutes of being discovered.

    Astra’s release is set to happen soon, though OpenAI has not given an exact date.


    Stop guessing and start investing with confidence. KnockoutStocks gives you the AI insights, market intelligence, and stock research you need to spot opportunities, cut through the noise, and make smarter investment decisions — all in one powerful platform.

    Sign up today and get 50% OFF full access to our premium stock picks.

    Simply use coupon code SPECIAL50 at checkout to claim your exclusive discount.



    Source link

    Stay in the Loop

    Get the daily email from CryptoNews that makes reading the news actually enjoyable. Join our mailing list to stay in the loop to stay informed, for free.

    Latest stories

    - Advertisement - spot_img

    You might also like...