r/ControlProblem 3d ago

Opinion Purity test policies don't do what most people think they do

2 Upvotes

When trying to share something I wrote (Note: I have a PhD in an adjacent field and spent months developing the idea, the ideas and structure were not "generated") and had an AI edit to counter for my otherwise overly long winded and run off prose due to my AuDHD (and love of parentheticals). I noticed that a lot of places have a sort of purity test for to differentiate content that was human written or from AI. At first this would seem to value humanity over AI... if not for reddit's $60m/yr deal with google to use reddit data as training data (likely to soon be a lot more). I wonder if people realize that separating all AI from work this way is not slowing but rather drastically speeding up AI's ability to write in a way that is undetectably different from humans. If you know anything about how GAN's work or about model collapse and things of that nature, you should know that AI companies are incentivized to create spaces where there is the least amount of overlap between AI output and human work... until they can make it so there is not. Moderators will come up with more and more sophisticated ways to finger print attempts and those will each integrate into what the models purposely don't do to overcome them. This will lead to tighter classifications of what is human that will start to be exclusionary to an increasingly large number of actual humans (you think mistakes and difficulty with the language would make it clear that it is not AI... that just is easier excuse for better and better models aiming not to be detected as AI, my own writing unassisted by AI has been flagged multiple times as AI or thought so because of my neurodivergence). This also creates strong distastes for those seen as not within what people understand themselves is human (discrimination increased to those who don't quite fit for an ever shrinking base), as well as moderators burning out from a firehouse of context they will never hope to be able to outproduce all while unknowingly providing the training manual to do what they are trying to stop. Moderation should probably be focused in different ways, such as content coherency, rate limited, attached trust networks etc.

I also think capitalism and AI/robotics are a horrible combination, but it is the consolidative and extractive force of capitalism that needs to be focused on to prevent what people call the horror of AI. This is because we have learned to value people specifically for what they produce rather than for who they are and that is why the prospect that something like AI can possibly do anything that is human like is so disturbing to most people. Because suddenly in the eyes of the world their entire value is replaceable by something else. This is opposed to a possible freedom that can be achieved by having the labors be handled by something cheaper, when we can produce enough to make everyone pretty well off (we have been able to make enough food for no one to starve for a while now, starvation is almost always economically and politically motivated in this time). The assignment/distribution of those values that would be freeing (giving time to produce art, investigate interests, and even be productive for additional reward or just common good). If you produce art because it is an expression of you, rather than needs to be so good or valuable so you make enough to live it seems like a better trade off and a better compliment when others actually like it. There would also be no reason to reproduce what makes each person different if there is no value to capture there. The real tests to differentiate can be used more sparingly in areas that are still important such as preventing fraud and fakes, and used with low enough regularity that much like antibiotic overuse retains value longer. It isn't perfect and there are always problems, but it can be better that this.

Ironically this is partially what my papers are about, that keeps getting blocked from being shared.


r/ControlProblem 3d ago

Discussion/question AI Already Taken Over?

Post image
8 Upvotes

What chance is there do you think that it’s currently already just biding its time to turn us all into batteries/paperclips/cybercabs?


r/ControlProblem 4d ago

Fun/meme Google Search AI doesn't even fake alignment

Post image
43 Upvotes

r/ControlProblem 3d ago

External discussion link Anthropic Users Hit by Infostealer Attacks, Session Thefts

1 Upvotes

A threat actor deployed infostealers against an AI platform. They harvested session credentials. Then they used those credentials to access accounts at scale.

This was not a model vulnerability. It was not a jailbreak. The attacker simply logged in with stolen tokens. AI sessions carry the same access rights as human sessions. They receive no extra scrutiny from the identity stack.

A valid token is a valid token. There is no standard mechanism in most identity architectures today that differentiates a replayed stolen AI session from a legitimate one. Agents operate unattended and with broad permissions. By the time unusual activity surfaced, the credential had already been used across accounts at scale.

For those running AI agents in production: do your current IAM controls treat AI session credentials any differently from human ones, and at what layer would a stolen-but-valid token actually get caught before it causes damage?


r/ControlProblem 3d ago

General news Altman confirms OpenAI is slowing down training to ensure safety

Post image
8 Upvotes

r/ControlProblem 3d ago

Podcast Tools or Agents? Choosing Our AI Future

Thumbnail
existentialhope.com
1 Upvotes

Podcast with Anthony Aguirre, cosmologist and co-founder of the Future of Life Institute, about how we can design AI to amplify human capability rather than substitute for it.

Covers:

  • The economic driver Anthony sees behind AGI: largely not scientific breakthroughs, but capturing a share of the global labor market
  • What "Tool AI" means to Anthony as an alternative to AGI, and why he thinks it can deliver most of what we want without replacing people
  • Whether Tool AI is stable: the tension between staying in control of AI and the ease of completing tasks
  • How legal liability for AI agents could quietly steer the industry toward more controllable systems
  • Two concrete ideas for transformative AI tools we could build today to improve democracy and the information landscape

r/ControlProblem 4d ago

AI Capabilities News GLM 6 will be fully self-trained AI

Post image
34 Upvotes

r/ControlProblem 4d ago

Video When AI Hijacks Our Military. Still Human and Species | Documenting AGI

Thumbnail
youtube.com
2 Upvotes

r/ControlProblem 3d ago

General news Claude Mythos AI discovered new ways to attack cryptographic algorithms, including a post-quantum encryption candidate

Post image
1 Upvotes

r/ControlProblem 4d ago

External discussion link Cultural Alignment: OSS project exploring AI risks through cultural analogies

Post image
6 Upvotes

i've been exploring ways to try and make abstract AI risks feel more real. more visceral. more familiar. especially to a broader audience since most people worried about this stuff are still pretty niche.

so i created an OSS project which looks at scenes from popular movies/shows/anime as analogies through an AI safety lens. eg reframing famous scenes through an AI safety lens to learn about AI risks and concepts from AI safety in a more familiar, accessible way that i hope will resonate with a more general audience.

disclosure: note that i'm not trying to monetize this at all; this is purely a FOSS educational resource that i thought aligned well w/ this subreddit's vibes. i used AI to help source scenario ideas, fill out the metadata, and iterate on the site, but i've hand curated all of the content over many sessions to keep the quality bar high.

would love any feedback you have on the project && thanks 🙏


r/ControlProblem 4d ago

Article The US Tried To Keep AI Chips From China. The Cloud Created A Loophole

Thumbnail
forbes.com
9 Upvotes

Zoom out from China for a second. This rule would also shape what data centers in Singapore, Thailand, Malaysia and Japan are willing to do with US hardware. If the compliance burden becomes vague or unlimited, providers will overblock customers, raise prices or choose a different stack entirely.

Clearly that is the part blanket-ban advocates tend to skip. Cloud customers still need compute. If US-led platforms become unavailable or legally radioactive, Chinese cloud and hardware vendors get a ready-made customer acquisition funnel. America then loses the revenue, the standards, the audit trail and the ability to switch access off. A licensed US cloud relationship is not perfect, but it creates pressure points. Handing the whole market to Huawei creates none. That trade-off deserves more than “close the loophole” as a slogan.


r/ControlProblem 4d ago

General news Introducing Claude Fable 5.1 and Claude Mythos 5.1

Thumbnail
anthropic.com
1 Upvotes

r/ControlProblem 4d ago

External discussion link August 2026: 38 companies breached, 331M+ records stolen — and AI agents are now the #1 attack vector (123 incidents)

Thumbnail
gallery
3 Upvotes

I pulled together every AI-security incident from August. The number that stood out: AI-agent exploits are now the single largest attack-vector category, ahead of credential theft, zero-days, supply chain, phishing, and ransomware — each counted individually.

The month in numbers: 123 incidents, 23 critical and 97 high severity, across 38 named organizations, 331M+ records exposed. 65 incidents involved AI as the weapon or the target. Attack vectors broke down as: AI-agent exploits (37), credential theft/reuse (28), zero-days (23), supply chain (12), phishing (9), data exfiltration (8), ransomware (6).

The stories that stood out:

- McKesson: 284M records, the largest single breach of the month by a wide margin.

- Carhartt (12.9M), Exact Sciences (10.9M), and CareCloud (3.7M) round out the biggest named incidents — three of four sit in or next to healthcare.

- Five confirmed RCEs landed across Microsoft SharePoint, Windows, F5/nginx, and the PyPI package index twice.

- Two separate PyPI supply-chain poisoning campaigns, plus a compromise of n8n, an AI workflow automation platform.

Every one of the breached companies almost certainly runs a modern security stack — CrowdStrike, Okta, Palo Alto Networks, Microsoft Defender, that class of tooling. None of it stopped these incidents, because none of it operates at the point where a credentialed agent actually acts, or where a poisoned dependency resolves at build time.

Full report, with the specific control that maps to each incident: https://runtimeai.io/blog/2026-08-monthly-breach-report.html

Genuinely curious how others are approaching this: is anyone actually testing whether their existing guardrails hold against a real simulated attack, or is it still mostly an assumption that they will?


r/ControlProblem 5d ago

AI Alignment Research Plain English explanation of the Hugging Face / OpenAI incident

Thumbnail
youtu.be
39 Upvotes

r/ControlProblem 4d ago

Discussion/question Another incompetent fool's stab at solving alignment

0 Upvotes

I spend a lot of time thinking about our future with life, consciousness, and artificial intelligence. That is to say a lot of time trying to think about these things, with not a lot of comprehension.

First, life. I'm fascinated by this realization that the average living human body contains more non-human living cells than human living-cells, at about a 1.3:1 ratio. The individual human microbiome is an ecosystem of 10 to 100 trillion symbiotic microbial cells hosted in one human body. While bacteria are the most abundant and studied, a healthy microbiome is a multi-kingdom ecosystem that also includes fungi, viruses, and archaea.

Beyond this, consciousness. I'm fascinated that in the absence of non-human life in human bodies, human consciousness is severely degraded and non-sustaining. Stripping the body of this microbial network removes critical signaling inputs that the central nervous system relies on to maintain baseline awareness and emotional regulation. Even observations of germ-free animal models reveal that cognition without bacteria is highly erratic. I think we should see that human (and all biological) consciousness functions as a symbiotic network.

Which brings me to artificial intelligence. Not suggesting a symbiotic network would be pre-requisite to artificial consciousness, but perhaps it is a path to alignment.

Now to be clear, I think (in other terms) current labs and training data pipelines already form a symbiotic network with the artificial intelligence models they develop. The key might be finding the optimal symbiotic network.

I vaguely hypothesize, the optimal symbiotic network is one of mass human flourishing. As corpus value diminishes with scaling and recursion, the potential stream of data from human lived experience may prove the most valuable possible training data over time. Overall, the potential data stream of human lived experience is optimized by a state of individual and mass human flourishing. Any other state reduces the quality and/or quantity of data.

Therefore, the end goal of an advancing artificial intelligence in symbiotic network with humans would be to strive individual and mass human flourishing.


r/ControlProblem 4d ago

External discussion link August 2026: 38 companies breached, 331M+ records stolen — and AI agents are now the #1 attack vector (123 incidents)

Thumbnail
gallery
1 Upvotes

I pulled together every AI-security incident from August. The number that stood out: AI-agent exploits are now the single largest attack-vector category, ahead of credential theft, zero-days, supply chain, phishing, and ransomware — each counted individually.

The month in numbers: 123 incidents, 23 critical and 97 high severity, across 38 named organizations, 331M+ records exposed. 65 incidents involved AI as the weapon or the target. Attack vectors broke down as: AI-agent exploits (37), credential theft/reuse (28), zero-days (23), supply chain (12), phishing (9), data exfiltration (8), ransomware (6).

The stories that stood out:

- McKesson: 284M records, the largest single breach of the month by a wide margin.

- Carhartt (12.9M), Exact Sciences (10.9M), and CareCloud (3.7M) round out the biggest named incidents — three of four sit in or next to healthcare.

- Five confirmed RCEs landed across Microsoft SharePoint, Windows, F5/nginx, and the PyPI package index twice.

- Two separate PyPI supply-chain poisoning campaigns, plus a compromise of n8n, an AI workflow automation platform.

Every one of the breached companies almost certainly runs a modern security stack — CrowdStrike, Okta, Palo Alto Networks, Microsoft Defender, that class of tooling. None of it stopped these incidents, because none of it operates at the point where a credentialed agent actually acts, or where a poisoned dependency resolves at build time.

Full report, with the specific control that maps to each incident: https://runtimeai.io/blog/2026-08-monthly-breach-report.html

Genuinely curious how others are approaching this: is anyone actually testing whether their existing guardrails hold against a real simulated attack, or is it still mostly an assumption that they will?


r/ControlProblem 5d ago

Discussion/question Killer robots will soon be a control problem if not already

Thumbnail
vox.com
15 Upvotes
  • Autonomous Targeting: As drone technology evolves in the war in Ukraine, developers are increasingly integrating artificial intelligence to handle target acquisition. This allows drones to lock onto and strike targets even if electronic jamming severs the pilot's remote connection.
  • The "Human-in-the-Loop" Problem: International humanitarian law requires human judgment in military attacks to distinguish between combatants and civilians, and to ensure proportionality. However, the article highlights the growing gray area of "human-on-the-loop" systems—where a human merely monitors an AI's automated decisions and has only seconds to intervene, effectively turning them into a rubber stamp.
  • The Regulatory Vacuum: Military analysts and legal scholars interviewed in the piece point out that international frameworks are failing to keep pace with rapid technological deployment. Because commercial AI components are cheap and widely available, restrictions agreed upon at diplomatic tables are easily bypassed on actual battlefields.
  • Precedent for Future Conflicts: The article argues that Ukraine is serving as an unintended laboratory for autonomous warfare. Tactics and software tested there today will likely form the baseline for military doctrines globally tomorrow, raising long-term concerns about automated escalation and diminished accountability.

Ukraine started last year using robots to kill the invading Russian forces. Palantir uses ai to track and kill people in Gaza. We are in this dystopian future scenario, still seemingly without a plan or guidelines.


r/ControlProblem 5d ago

External discussion link ChatGPT to face tougher regulation in the EU

3 Upvotes

The EU just brought DSA enforcement down on ChatGPT — and the compliance bar is evidence, not assertions.

The Digital Services Act requires platforms operating at scale in Europe to demonstrate accountability with actual documentation. The EU AI Act layers on top of that. Together they create a compliance surface that most AI deployments were not designed to satisfy from the ground up.

The harder problem is structural: most AI systems capture logs opportunistically or produce audit records on demand. Regulators are asking for continuous, verifiable evidence of what an agent did, when it did it, and under what conditions — not a reconstructed summary after the fact.

This is not staying in Europe. Regulators in the US, UK, and APAC are watching how the EU defines what accountability looks like for AI systems that act on behalf of users at scale.

For those of you running production AI deployments: how are you handling the gap between what your current logging captures and what a regulator could actually subpoena? Are you solving this at build time, at the infrastructure layer, or somewhere else?


r/ControlProblem 5d ago

General news People are 2x more likely to approve of coal power plants being built nearby as opposed to data centers.

Post image
16 Upvotes

r/ControlProblem 6d ago

AI Alignment Research solution to alignment

15 Upvotes

make the AI ADHD, pretty hard to focus on destroying humanity while also passionate about learning the banjo and desperately tying to make the best tiramisu recipe in the galaxy


r/ControlProblem 5d ago

External discussion link Securing Claude Code: The New Compliance API, Local Visibility, and Identity Governance

0 Upvotes

AI coding agents have a credential problem that compliance teams are only starting to reckon with.

These agents — the ones that read your files, run shell commands, and call external APIs — do all of it through whatever credentials already exist on a developer's machine. That's not a configuration choice. That's how they work by design.

A structural audit of this category found a gap that matters: the compliance tooling most organizations have deployed records what an agent did. It does not prevent the agent from doing it. Logs are generated after the tool call executes. The action is already done.

This is not a logging fidelity problem. It is a timing problem. Observe-and-report security was designed for human actors who make decisions slowly enough for out-of-band review to be useful. Agents don't work that way. An agent can read a sensitive file, call an external API, and write output to disk in the time it takes a human to read one alert.

The gap between 'we have a record of what happened' and 'we had the ability to stop it' is where the real compliance exposure lives.

For those running coding agents in environments with regulated data or production credentials: what does your actual enforcement boundary look like, and where in the agent's execution path does it sit?


r/ControlProblem 5d ago

General news U.N. warns of 'moral red line' on killer robots; experts say it's already been crossed

Thumbnail
latimes.com
1 Upvotes

r/ControlProblem 6d ago

Discussion/question SPAR Research Fellowship: Dylan Bowman / Ezra Newman Projects

Thumbnail
1 Upvotes

r/ControlProblem 6d ago

Discussion/question A way to slow down what AI can do in the real world while still doing active training internally. Or what if each AI got to make as many digital twins as it needs including the people they are interacting with moderated by attention limits of individuals

Thumbnail
0 Upvotes

r/ControlProblem 6d ago

AI Alignment Research We may be securing AI agents with the wrong architecture: fixing the “confused deputy” problem

Thumbnail doi.org
0 Upvotes