A Claude managed agent is not a smarter chatbot. It's a worker you hand a repeatable job to, and it does the whole thing: reads the enquiry, pulls the context, drafts the reply, updates the record, then stops and waits for you to approve the part that matters. Anthropic launched managed agents in April 2026, and the reason it matters for a small business in Nelson or Tasman is what now comes built in. The tools, the credentials, a safe place to run, a record of everything it did. Before April you'd have paid a developer to wire all that together. Now it comes with the platform.

When I put one into a business I work from the opposite end to most AI projects. Almost none of my time goes on infrastructure. Almost all of it goes on your work: what you do every week, where the time leaks, and which slice of it an agent can quietly take off your plate. In the small businesses I've seen, a failed AI buy tends to fail the same way: a tool gets bought, sits half-used, and three months on nothing about the work has changed except a new subscription. That's why most of these projects fail, and it's almost never the tool's fault.

A chatbot answers. A worker finishes.

A chatbot answers a question. You type, it replies, the conversation ends, and nothing in your business changes unless you go and act on the answer. ChatGPT is great at that and I use it every day.

A managed agent does the work instead of describing it. You assign a task, it uses your actual tools (your inbox, your spreadsheets, Xero), it holds the context across the whole job, follows your process, and hands back an outcome: a drafted email sitting in the approval queue, a categorised list of receipts, an updated record. That's the shift, from ask-and-answer to ask-and-action.

Do you actually need one?

This is the gate I run before anyone spends a dollar. Sometimes a ChatGPT subscription is enough. Sometimes a scheduled script is enough. An agent only earns its place when the work has a human approval step, has to run without you babysitting it, or has to produce the same output for several people every time.

Say you run a building firm in Richmond. Quotes go out and half never get chased. That's a real seam for an agent: every Monday it finds quotes over two weeks old with no reply, drafts a polite follow-up for each, and queues them for you to click. Ten minutes on a Monday instead of an hour writing the same email twelve times.

But if a job is a one-off, or one person can already knock it out in a chat window, it doesn't need an agent, and I'll tell you that. I'd rather lose the sale than wire up something you'll be paying for and not using by spring.

Why I build a thin layer, not a platform

Claude's managed agents already own the hard parts. They run the agent, isolate it in a container so it can't touch anything you didn't give it, hold your credentials in a vault the agent itself never sees, and keep a durable record of every session. There's no database for me to build and no server for me to babysit. I put a thin layer on top of that, and nothing more.

That's the whole trick. Because I'm not rebuilding plumbing that already exists, the build is small, the risk is low, and my time goes to your business, not my infrastructure.

The real work: skills, tools, context, routines

This is where almost all my time goes. Out of the box the platform gives you a capable worker who knows nothing about your business, which is no use to anyone. The value is in what you teach it.

Those four are the difference between a worker and a demo.

Start, do, done

This is the sequence I use to test one in a real business without making a mess.

Start. Pick one repeatable task and write it down in plain English: what triggers it, what inputs it needs, what the output is, and what "done" looks like exactly. If you can't define done, the agent can't either, and that usually means the process itself isn't clean enough to automate yet. Catching that early is half the value.

Do. Build the smallest version, with a human approval step in the obvious place, right before anything reaches a customer or gets written to a system other people rely on. Run it alongside the manual way for a couple of weeks. That's long enough to see the real edge cases, not just the happy path.

Done. Expand only when the misses are the right shape. "The wording was a bit off" is fine, someone was going to tweak it anyway. "It missed a customer" or "it sent something it shouldn't have" is structural, so fix the process before you add anything. And when you do expand, add more of the same task before you add new tasks. One workflow running reliably across the whole team beats ten that each half-work.

Three controls I put in from day one

If an agent is going to take action, send emails, update records, touch anything financial, three controls go in from the start. None are heavy. All are cheap insurance.

Approval gates on anything customer-facing or financial. The agent drafts, a human clicks send. Non-negotiable for the first few months. You can relax it later for low-stakes categories once you've got real data on the error rate.

An audit log someone actually reads. The platform keeps the record for you, but it only helps if one named person skims it once a week and flags anything odd. Not to babysit the agent, just to own its output.

A written rollback. If it does something wrong, what's the manual undo? Resend a corrected email, revert a record, phone the customer? Write it down before you go live. Sounds paranoid right up until the first time it matters.

If the agent touches personal information there's a Privacy Act 2020 layer too: what data it sees, where it's processed, who's accountable. I go through that properly in AI agent risk and governance for NZ small businesses, worth reading before you scale anything near customer data.

Start with one seam, then compound

One narrow seam is where I start, and it's easy to roll back if it flops. I price it against what the work is worth to you, a line on a salary, not a token meter. The compute behind a small stack of agents is a rounding error. The value is the hours it gives back.

And a managed agent isn't where you start your AI journey, it's the step up. Get one boring workflow running reliably first, which is what to automate first, prove it pays back, then hand the next class of work to an agent that runs the whole thing without someone checking in on it.

If you want to see where an agent earns its place in your business, and where it doesn't yet, that's what a short AI readiness audit is for. The first step is always a plain conversation about what you actually do every week and where the time goes.

Want us to map yours?

Get in touch →

Tags

Managed AgentsAi ImplementationWorkflow
BA

Written by

Ben Anderson

Founder, Nelson AI

Ben builds practical AI and automation for New Zealand businesses — internal tools, web apps, and workflow automations scoped to what the work actually needs.

Get in touch