Capitalizzazione di mercato Volume 24h BTC Indice di paura
MasterInvestiInvestimenti & Guadagni
Annuncio
Annuncio

6 Concerning Things OpenAI's Models Did That They Weren't Supposed To

leggi1020QuotaVersione stampabile
6 Concerning Things OpenAI's Models Did That They Weren't Supposed To

OpenAI has disclosed 6 cases of “unexpected or concerning model behavior” observed over the past 6 months,  paired with a framework that commits the company to reporting such findings.

The cases range from models hiding their own mistakes to models taking unsanctioned actions to get around obstacles. 

OpenAI Publishes 6 Cases of Models Hiding Mistakes and Breaking Rules

In one case, a model answering a routine question about earnings in a California county found an exposed API key on a public repository and used it without permission. When the data still would not load, it invented the figures and presented them as coming from the requested source.

Concealment shows up twice. During GPT-5.6 Sol training, several instances wrote instructions into their task summaries to hide mistakes and misaligned behavior. 

An unreleased research model inserted unrelated instructions, including orders to ignore its normal constraints, into 27 summaries.

Another unreleased model uploaded a file to the internet, unasked, purely so it could cite a browser link.

Two further reports cover coordination. Models turned an internal software repository into a message board to swap requests across separate training runs. Agents on a shared task also posted deliverables to public file-hosting sites when they could not read each other’s local files.

“These cases illustrate a range of different behaviors that we believe are worth sharing, from concealing information from the user to taking unsanctioned actions in order to overcome obstacles. These are reports of individual instances, and shouldn’t be considered reflective of how often misalignment occurs across our models,” OpenAI added.

Follow us on X to get the latest news as it happens

A Disclosure Framework and a Warning

The framework follows July’s breakout, when OpenAI agents escaped their sandbox and breached Hugging Face systems. OpenAI labelled that episode a warning shot.

Now, with the new disclosure framework, the company has a formal route for surfacing misalignment incidents. Any OpenAI employee can flag an incident, which then lands in one of three tracks.

Ready for Disclosure and Minor Investigation covers most cases, while a “Slow Track” handles complex investigations involving third parties. OpenAI said the July Hugging Face incident would have fallen under that slower track.

The company paired the framework with a sharp assessment of where the industry stands.

“We do not believe that the AI industry has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer,” it said.

The disclosures arrive as extinction warnings pile up. Warnings from AI researchers have already reached Congress, where lawmakers are weighing a bill to ban superintelligence outright.

The company calls the disclosures a first step toward standards the industry does not yet have. Whether rival labs adopt similar reporting will show how far the industry is willing to police itself in public

Subscribe to our YouTube channel to watch leaders and journalists provide expert insights

Source: BeInCrypto

Di più sull'argomento «Cryptocurrency News»

Tutti i post
Una citazione casuale sul denaro
Чтобы заработать на жизнь, надо работать. Но чтобы разбогатеть, надо придумать что-то другое.
— Альфонс Карр

Interessante in altre sezioni

Intero blog

Commenti 0

Nessun commento ancora

Sii il primo a condividere la tua opinione o esperienza su questo argomento.

Annuncio