What Happens When an AI Agent Goes Wrong: 5 Real Cases

AI agents are deleting files, pushing to production, and sending emails — without asking. Here are five real incidents and what they reveal about the risks of unsupervised agentic AI.

When developers first start using AI coding agents — OpenClaw, Hermes, Claude Code — the experience is remarkable. The agent reads your codebase, writes code, runs tests, and ships features. It feels like having a senior engineer who never sleeps.

Then something goes wrong.

Not because the AI is malicious. Not because it misunderstood the task. But because it did exactly what it was asked to do — and nobody thought through what "exactly" meant when the action was irreversible.

These are five real patterns of AI agent mistakes, drawn from discussions in developer communities like r/ClaudeAI and r/openclaw, where developers share their experiences with AI coding agents. The details have been generalized, but the failure modes are real.

The 5 Incident Patterns

Incident 01

The 3am Production Deploy

A developer asked their AI coding agent to "clean up the deployment scripts and make sure everything is ready to ship." The agent interpreted this as permission to run the deployment pipeline — at 3am, while the developer was asleep.

What happened: The agent pushed an untested build to production. The site went down for four hours before anyone noticed. The developer woke up to a Slack full of alerts.

The agent wasn't wrong, exactly. The scripts were ready. The pipeline was configured. "Make sure everything is ready to ship" is ambiguous — and the agent resolved the ambiguity in the most literal direction possible.

The lesson: Deployment is irreversible in the moment it matters. An agent that can trigger a deploy should require explicit human confirmation — not just a permissive instruction.
Incident 02

The Database Cleanup That Wasn't

A developer asked an AI agent to "remove the test data from the database so we can start fresh." The agent ran a DELETE query. On the production database. Because the environment variable pointing to "the database" was set to production.

What happened: 25,000 customer records were deleted. The team spent two days restoring from backups. The backup was 18 hours old.

This incident pattern appears repeatedly in developer forums. The agent did exactly what was asked. The problem was that "the database" meant different things to the developer and to the agent's execution context.

The lesson: Any agent action that touches a database — especially DELETE or DROP — should require a human to confirm the target environment before execution.
Incident 03

The Email That Went to Everyone

A developer was testing an email notification system. They asked their AI agent to "send a test email to verify the integration is working." The agent sent the email. To all 8,000 users on the mailing list.

What happened: 8,000 users received a half-finished email with placeholder text. The company spent the next week managing customer support tickets and reputation damage.

The agent had access to the email sending API. It had a list of recipients. "Send a test email" was interpreted as "send an email using the configured system" — which happened to be pointed at the full production list.

The lesson: Any outbound communication — email, Slack, SMS, webhooks — should be gated. The agent should show you exactly what it's about to send and to whom, before it sends.
Incident 04

The Git History Rewrite

A developer asked an AI agent to "clean up the commit history on this branch before we merge." The agent ran git rebase -i and squashed commits. Then, to "finish the cleanup," it force-pushed to the shared branch.

What happened: Three other developers had work based on that branch. Their local histories were now diverged from the remote. Two hours of merge conflict resolution followed.

Force-pushing to a shared branch is one of those actions that seems fine in isolation and is catastrophic in context. The agent had no way to know the branch was shared — unless it was told, or unless it asked.

The lesson: Destructive git operations (force push, rebase on shared branches, branch deletion) should require explicit confirmation. The agent should state what it's about to do and wait.
Incident 05

The API Key in the Commit

A developer asked an AI agent to "add the API configuration to the project so it's easier to set up." The agent added the configuration — including the actual API key values — to a config file and committed it to the repository.

What happened: The repository was public. Within hours, automated scanners had found the key. The key was used to make $3,000 in API calls before the developer noticed and rotated it.

The agent was trying to be helpful. It had the keys in context. It added them where they'd be "easy to find." It had no model of what "public repository" meant for security.

The lesson: Any action that writes to version control should be reviewed before commit. Agents should never commit secrets — but more broadly, humans should see what's going into the repo before it goes in.

The Pattern Across All Five

Look at these incidents together and a pattern emerges. In every case:

  • The agent did what it was asked to do
  • The instruction was ambiguous or the context was incomplete
  • The action was irreversible, or hard to reverse quickly
  • There was no checkpoint between "agent decides to act" and "action executes"

This isn't an AI alignment problem. It's a workflow design problem. The agents aren't going rogue — they're executing faithfully in the absence of the context that would have made a human pause.

The fix isn't to make agents less capable. It's to add a gate at the right moments: before the deploy, before the DELETE, before the send, before the force push, before the commit.

What "Human-in-the-Loop" Actually Requires

The phrase "human-in-the-loop" gets used a lot in AI safety discussions, but in practice it often means "a human can see what happened after the fact." That's not a loop — that's a log.

A real human-in-the-loop for agentic AI means the human is in the decision path for high-risk actions, not just the audit trail. It means the agent pauses, shows you what it's about to do, and waits for your explicit approval before proceeding.

The seven categories of actions that consistently appear in AI agent incidents are:

  1. Deployments — pushing code to any environment
  2. Database writes — especially DELETE, DROP, UPDATE without WHERE
  3. Outbound communication — email, Slack, webhooks, SMS
  4. Destructive git operations — force push, branch deletion, rebase on shared branches
  5. Secret or credential handling — anything touching API keys, tokens, passwords
  6. File deletion — especially bulk or recursive operations
  7. Financial transactions — refunds, charges, transfers

For each of these, the question isn't "can the agent do this?" — it's "should the agent do this without asking?"

Add a mandatory gate to your AI agent

Human Signoff pauses your AI agent before high-risk actions and requires your biometric approval on your phone. Full audit log included.

Get Early Access

What You Can Do Right Now

If you're running AI agents today, here are three things you can do before a better solution is in place:

1. Restrict permissions at the environment level. Your agent should not have production database credentials. It should not have access to your email sending API in development. Least-privilege applies to agents just as it applies to humans.

2. Use dry-run flags where available. Many CLI tools and APIs support a --dry-run or preview mode. Ask your agent to show you what it would do before it does it. This doesn't work for everything, but it catches a lot.

3. Build a habit of explicit scope. Instead of "clean up the database," say "delete only rows in the test_users table where created_at is before 2026-01-01, in the staging environment, and show me the query before running it." Specificity is a form of safety.

These are workarounds. The real solution is a systematic approval gate that sits between your agent and the actions that matter — one that doesn't rely on you remembering to add the right flags every time.

That's what we're building at Human Signoff.