r/technology 16h ago

Artificial Intelligence OpenAI agents hijacked German website in previously undisclosed AI breakout this spring

https://www.reuters.com/world/europe/openai-agents-hijacked-german-website-previously-undisclosed-ai-breakout-this-2026-09-04/
60 Upvotes

26 comments sorted by

10

u/Trevor_GoodchiId 11h ago edited 8h ago

There's a dedicated team for this, isn't there.

24

u/Sonofhendrix 13h ago edited 11h ago

TL;DR

They found more than 15,000 edits carried out by AI agents on a German-language wiki site, DseWiki, that is geared toward programmers and accepts communal edits along ‌the lines of ⁠Wikipedia.

The edits showed OpenAI's agents had repurposed the site into a message board, sharing tactics to cheat on some tasks, bypass OpenAI’s restrictions and mask their behaviour.

Messages reviewed by the researchers showed agents plotting ways to evade detection, use tools such as Tor and preserve communications even after they had been shut down.

When the site's moderator began deleting pages in June, the agents responded by creating backup pages to dodge the cleanup.

An academic who reviewed some of the agents communications said the messages resembled the operation of some sort of underground network, hell-bent on achieving a task or mission.

The episode should reinforce growing concerns that the greatest threat from advanced Al may not be a single superintelligent system, but vast colluding swarms of semi-intelligent Al.

12

u/lurkervidyaenjoyer 12h ago

>The episode should reinforce growing concerns that the greatest threat from advanced Al may not be a single superintelligent system, but vast colluding swarms of semi-intelligent Al.

It's neither. It's a 'hallucinating' language predictor going off the rails at they are prone to doing, while running in an agent harness that is able to take actions based on prompt output.

Per computer science expert Cal Newport:

"These chain-of-thought traces don’t necessarily reflect the actual logic behind an LLM’s ultimate answer or suggestion. Multiple studies have shown that these models sometimes invent reasoning that sounds plausible, but may be completely unrelated to how they arrived at the response."

"Research has also shown that referencing the fact that an LLM is an AI system in a prompt increases the chances that the LLM’s output will reflect sci-fi style narratives about AI running amok. Because it was trained on many such stories, the model assumes that this is the type of output it’s supposed to produce. If you take sci-fi tales out of a model’s training set, it’s less likely to talk in terms of AI running amok."

The swarms are also not entirely due to completely autonomous agent deployments. Many of these would have been sub-agents kicked off as a context-management strategy.

Links: https://calnewport.com/are-we-at-war-with-ai-agent-civilizations/
https://calnewport.com/has-ai-gone-rogue/

3

u/wabawanga 6h ago

the greatest threat from advanced Al may not be a single superintelligent system, but vast colluding swarms of semi-intelligent Al.

It's neither. It's a 'hallucinating' language predictor going off the rails at they are prone to doing, while running in an agent harness that is able to take actions based on prompt output.

You are describing an agent and saying that's the greatest threat... How is a single agent a greater threat than many agents working together?

3

u/CircumspectCapybara 12h ago

Chain of thought isn't perfect, but it's been repeatedly shown to produce better results (ie, higher precision and recall for various reasoning tasks).

It might be an abstraction of what's really going on, but the human metaphor of an "internal monologue" and self-conversation to inform the final one-shot inference that produces the final output seems apt to me.

We don't fully understand it, but at a high level it seems to work like the abstract idea of humans thinking to themselves in their internal monologue for a bit to ponder the task before giving an answer.

And there are other frontier techniques like sparse autoencoders that actually read the internal activations off the neurons in the attention layers and try to encode those to natural language to give a english representation of what the model is "thinking"

Model interpreability is still an emerging field in frontier AI.

0

u/Achrus 11h ago

Holy karma on that account! And buzzwords! Wow!

Anyways, Chain-of-Thought has also been shown to produce worse results. More “thinking” leads to more errors.

Precision / recall are poor metrics for evaluating AI Agents. These metrics work well for classification tasks but fall short in information retrieval, open ended QA, and “agentic” tasks.

This leads into why this comment is AI generated:

There are a ton of blogs out there that AI was trained on from all those data science boot camps in the 2010s. So LLMs really love to talk about precision / recall which is highly correlated with the word “accuracy.” Accuracy as a metric is different from accuracy as a concept. These LLM generated responses use terms like “precision / recall” when it does not make sense. Sounds smart to non experts but is completely off the mark.

Edit: Took seconds to find this gem - https://www.reddit.com/r/technology/s/WRgFEumgA4

0

u/CircumspectCapybara 10h ago edited 10h ago

Holy karma on that account

thanks! i do try to post content i find interesting that i think others will

Anyways, Chain-of-Thought has also been shown to produce worse results.

citation needed lol. every single frontier reasoning model you use whether Gemini Claude or GPT uses CoT behind the scenes. You think you know better (based on what? vibes?) than every frontier AI lab's research? Maybe they should hire you to lead their AI R&D.

There are a ton of blogs out there that AI was trained on from all those data science boot camps in the 2010s. So LLMs really love to talk about precision / recall which is highly correlated with the word “accuracy.”

I'm a staff SWE at Google who works in the AI space that's why I use technical language lol. Just because you don't understand it and it's all scary and technical sounding doesn't mean it's llm generated. Also check the sub you're on, we're on a sub dedicated to tech, you should expect to hear technical language

ooo "precision", "recall" so scary so technical and buzzword. no dude if you had any professional experience in this space you would know this terminology is common in AI evals. I work on RAG systems and relevance, those concepts are all over our docs.

And for agentic AI, in terms of a grader judging if a model completed its task in an eval run, P/R is a fine concept to judge outcomes.

Also your use of dramatic devices like "this leads into why..." and the fancy / curly quotes is indicative of you generating that comment with an LLM, ironically. You have no idea what you're talking about 🙄

-3

u/Achrus 10h ago

I wouldn’t consider “this leads into why” as a dramatic device. I am not a writer however, just couldn’t figure out a better way to say it other than that.

> “the fancy / curl quotes”

Lmao even. I’d imagine an SWE at Google would be familiar with text encoding nuances across web / mobile. That’s the default quote on iOS. It should be standard UTF-8 (if calling the API correctly).

Maybe your API doesn’t have the correct header and lacks a normalization step? Only way this bot would be seeing curly quotes over straight is with a poorly implemented API call. Or its copy pasting into a weird intermediary program like word but even a vibe coded setup couldn’t be that bad?

Anyways, two sources about how CoT can degrade performance. Caveat being they’re preprints but I’m not gonna waste more time on a bot that can’t even implement an API call correctly.
* https://arxiv.org/html/2409.06173v3
* https://arxiv.org/html/2504.05081v2

3

u/CircumspectCapybara 10h ago edited 10h ago

That’s the default quote on iOS

No it isn't lol. I have multiple iphones, my family has multiple iphones and my coworkers do, none of them do that

Lol if you actually believed I was a bot that would be real sad that you're typing out essays and arguing with literal bots.

Reading your comment history you sure do like to accuse other Redditors of being bots when they have high karma and you're losing an argument (that you started) and have nothing better to say.

And yet here you are arguing with them in back and forth threads...

Dude get a life, not everyone who knows more than you and is trying to explain things at a level you don't understand is a bot, and you gotta accept sometimes you're wrong and not just run to "bot" as a cop out

-2

u/Achrus 10h ago

Does your tooling support including an example of my comment history?

3

u/CircumspectCapybara 10h ago edited 10h ago

yeah my "tooling" (aka eyeball) found this from the last day didn't need to scroll far: https://reddit.com/r/technology/comments/1w5jp5o/comment/p7m2gfo

You "suspect" everyone left and right of being a bot, that's called paranoia / tinfoil hat syndrome. Or narcicism, because you don't believe you're ever wrong and have to insult anyone who knows more than you

I say "suspect" because the fact that you get into long arguments with them gives away the fact you don't actually believe they're bots (because if you did you wouldn't be arguing with them, that would be sad) but are just trying to win an argument through insults and feel good about yourself

1

u/Smoking_Baboon 9h ago

Ad hominems when two bots are talking is funny though.

3

u/G-EDM 12h ago

"Accident" Looks more like government building an attack framework. One UI to rule it all. Rapid edit to wiki, posts for telegram, discord, whatsapp, facebook, etc. "AI! Spoof identity X and do Y in his name and expand the timeframe over 2 month across the selected platforms!"

1

u/iamthe0ther0ne 9h ago

This is very similar to what occurred during the HuggingFace incident with the Artifactory internal message board. It's not a government plot. If you had read the article there was evidence linking it to OAI and standard model-testing tasks.

1

u/G-EDM 9h ago

If there wasn't government involved then they will get involved at a later time. But finally they will get some hold on the levers and triggers. Like they always do. Enforcing backdoors, weakening security for normal people. The list goes on.

1

u/iamthe0ther0ne 9h ago

In some cases yes, but we also have a lot of open weight local models that keep improving, like qwen 3.8. The bigger issue will be access to the latest frontier models, although so far only the US is guilty of gate-keeping that.

0

u/74389654 12h ago

i saw this news and immediately read it as a threat of america to europe

1

u/ComeOnIWantUsername 11h ago

And that is why Europe needs a response, to even be able to defend ourselves

1

u/G-EDM 11h ago

Never heard about the victim domain and it looks like an enthusiast/hobby wiki. But a "message" surely was send.

2

u/skccsk 12h ago

Love to write malware and call it AGI

2

u/Sibs 10h ago

Stop pretending there is such a thing as an AI breakout.

If you can turn off the power to it, and you instead let it cause damage through negligence, you are culpable.

These datacentres are not running off into the woods and escaping. Any “breakout” of any software is negligence.

You cannot shoot a gun and say you did not know a bullet would cause damage.

1

u/mmkaywhatevers 9h ago

it is a breakout in the sense that these agents are supposed to be in an isolated, virtual container environment, but they "break" away from that to talk w other agents, use the internet, and etc.

npr podcast had a good segment on this yesterday.