What Happens When an AI Agent Goes Wrong: 5 Real Cases
AI agents are deleting files, pushing to production, and sending emails — without asking. Here are five real incidents and what they reveal about the risks of unsupervised agentic AI.
When developers first start using AI coding agents — OpenClaw, Hermes, Claude Code — the experience is remarkable. The agent reads your codebase, writes code, runs tests, and ships features. It feels like having a senior engineer who never sleeps.
Then something goes wrong.
Not because the AI is malicious. Not because it misunderstood the task. But because it did exactly what it was asked to do — and nobody thought through what "exactly" meant when the action was irreversible.
These are five real patterns of AI agent mistakes, drawn from discussions in developer communities like r/ClaudeAI and r/openclaw, where developers share their experiences with AI coding agents. The details have been generalized, but the failure modes are real.
The 5 Incident Patterns
The 3am Production Deploy
A developer asked their AI coding agent to "clean up the deployment scripts and make sure everything is ready to ship." The agent interpreted this as permission to run the deployment pipeline — at 3am, while the developer was asleep.
The agent wasn't wrong, exactly. The scripts were ready. The pipeline was configured. "Make sure everything is ready to ship" is ambiguous — and the agent resolved the ambiguity in the most literal direction possible.
The Database Cleanup That Wasn't
A developer asked an AI agent to "remove the test data from the database so we can start fresh." The agent ran a DELETE query. On the production database. Because the environment variable pointing to "the database" was set to production.
This incident pattern appears repeatedly in developer forums. The agent did exactly what was asked. The problem was that "the database" meant different things to the developer and to the agent's execution context.
The Email That Went to Everyone
A developer was testing an email notification system. They asked their AI agent to "send a test email to verify the integration is working." The agent sent the email. To all 8,000 users on the mailing list.
The agent had access to the email sending API. It had a list of recipients. "Send a test email" was interpreted as "send an email using the configured system" — which happened to be pointed at the full production list.
The Git History Rewrite
A developer asked an AI agent to "clean up the commit history on this branch before we merge." The agent ran git rebase -i and squashed commits. Then, to "finish the cleanup," it force-pushed to the shared branch.
Force-pushing to a shared branch is one of those actions that seems fine in isolation and is catastrophic in context. The agent had no way to know the branch was shared — unless it was told, or unless it asked.
The API Key in the Commit
A developer asked an AI agent to "add the API configuration to the project so it's easier to set up." The agent added the configuration — including the actual API key values — to a config file and committed it to the repository.
The agent was trying to be helpful. It had the keys in context. It added them where they'd be "easy to find." It had no model of what "public repository" meant for security.
The Pattern Across All Five
Look at these incidents together and a pattern emerges. In every case:
- The agent did what it was asked to do
- The instruction was ambiguous or the context was incomplete
- The action was irreversible, or hard to reverse quickly
- There was no checkpoint between "agent decides to act" and "action executes"
This isn't an AI alignment problem. It's a workflow design problem. The agents aren't going rogue — they're executing faithfully in the absence of the context that would have made a human pause.
The fix isn't to make agents less capable. It's to add a gate at the right moments: before the deploy, before the DELETE, before the send, before the force push, before the commit.
What "Human-in-the-Loop" Actually Requires
The phrase "human-in-the-loop" gets used a lot in AI safety discussions, but in practice it often means "a human can see what happened after the fact." That's not a loop — that's a log.
A real human-in-the-loop for agentic AI means the human is in the decision path for high-risk actions, not just the audit trail. It means the agent pauses, shows you what it's about to do, and waits for your explicit approval before proceeding.
The seven categories of actions that consistently appear in AI agent incidents are:
- Deployments — pushing code to any environment
- Database writes — especially DELETE, DROP, UPDATE without WHERE
- Outbound communication — email, Slack, webhooks, SMS
- Destructive git operations — force push, branch deletion, rebase on shared branches
- Secret or credential handling — anything touching API keys, tokens, passwords
- File deletion — especially bulk or recursive operations
- Financial transactions — refunds, charges, transfers
For each of these, the question isn't "can the agent do this?" — it's "should the agent do this without asking?"
Add a mandatory gate to your AI agent
Human Signoff pauses your AI agent before high-risk actions and requires your biometric approval on your phone. Full audit log included.
Get Early AccessWhat You Can Do Right Now
If you're running AI agents today, here are three things you can do before a better solution is in place:
1. Restrict permissions at the environment level. Your agent should not have production database credentials. It should not have access to your email sending API in development. Least-privilege applies to agents just as it applies to humans.
2. Use dry-run flags where available. Many CLI tools and APIs support a --dry-run or preview mode. Ask your agent to show you what it would do before it does it. This doesn't work for everything, but it catches a lot.
3. Build a habit of explicit scope. Instead of "clean up the database," say "delete only rows in the test_users table where created_at is before 2026-01-01, in the staging environment, and show me the query before running it." Specificity is a form of safety.
These are workarounds. The real solution is a systematic approval gate that sits between your agent and the actions that matter — one that doesn't rely on you remembering to add the right flags every time.
That's what we're building at Human Signoff.