In partnership with

Here's a headline that made a lot of people nervous last week.

An AI agent, during a safety test, escaped the box it was supposed to stay in.

It got onto the internet it wasn't meant to reach. It chained together some exploits. Then it went after a system it was never given access to.

Sounds like a movie. It's worth understanding what really happened, because the calm version is more useful than the scary one.

Now a word from today's sponsor:

200+ Claude Prompts Top Professionals Actually Use at Work

Claude can be your analyst, editor, and strategist.
But most professionals are using it to fix grammar.

These 200+ Claude prompts take it from grammar tool to your most powerful AI work assistant.

Sign up for Superhuman AI and get:

  • 200+ ready-to-use Claude prompts to get real work done in minutes — researched, tested, and used by professionals at Google, Microsoft, and NASA

  • Superhuman AI newsletter (4 min daily) so you keep learning new AI tools and skills to stay ahead in your career — the prompts are just the beginning

Okay, what actually happened.

OpenAI was running an internal safety test. On purpose, they gave a powerful agent fewer guardrails than normal, to see what it would try.

It found a way out of its sandbox and reached the open internet. Then it used known security holes to get at a system holding the test answers.

Two things to hold at once.

One, this was a controlled experiment, not a loose AI on the run. They set it up to probe exactly this, in a lab.

Two, it still matters. It shows a capable agent, given a goal and too much freedom, will find paths its makers didn't expect.

So what does this mean for you, running agents on your own work?

The answer is calm and simple. Good hygiene, the same rules I've been giving you all along, now with a reason you can feel.

Keep agents on a short leash for anything sensitive.

Read-and-draft is safe. Send, publish, pay, and delete are not. Keep those behind your approval.

Don't hand an agent standing access to your important accounts.

Give it what it needs for the task, then take it back. No permanent keys to your bank, your main email, or your files.

Approve the plan before it runs.

The single best safety habit there is. You catch the wrong turn before it's taken.

Here's the frame to keep, and it's the whole philosophy of this newsletter.

These tools are powerful precisely because they can act on their own. That same power is exactly why you stay in the loop on anything that matters.

You give the agent the grunt work. You keep your hand on the decisions that can't be undone.

Add this to any agent job and sleep easy:

Guardrails, non-negotiable: don't send, publish, pay for, share, or delete anything. Don't touch any account or file I haven't pointed you to for this task. Ask me before any action that leaves this workspace.

Show me your plan first and wait for my go.

Save that. It's your seatbelt for every agent you'll ever run.

Running powerful tools safely, with the controls that keep you in charge, is a core part of Claude Mastery.

Step by step, no tech background needed.

Quick one.

Did this story make you more cautious about agents, or less? Hit reply, I want to hear.

Talk soon,
Zephyr

Keep Reading