What to Let an Agent Do While You Sleep

There are two standard answers to the question of what an agent should be allowed to do while you are asleep, and both of them are answers to a different question. The first is nothing, obviously — it drafts, you approve, that is the whole arrangement. The second is whatever it is good at — if it classifies tickets correctly ninety-six times out of a hundred, let it classify tickets. One policy gives up all the value. The other measures the wrong thing entirely.
Capability is not the axis that matters. A model that is 96% accurate at tagging and 96% accurate at drafting refund replies is the same model, and those two jobs are nowhere near the same risk. Four wrong tags out of a hundred is a slightly noisy tag cloud you fix with one sweep on Tuesday. Four wrong refund replies out of a hundred is four people who now believe something untrue about their own money, sent from your address, with no unsend button.
So the useful question is the one an insurer asks, not the one a benchmark asks: what does a mistake cost, and who finds out first? Sort overnight work along that axis and it falls into three tiers — plus one category that looks like the safest tier and belongs in the most dangerous one. That last case is where most people get burned, so it gets its own section.
Tier one: being wrong is cheap, and you find out first
The first tier is everything where the worst realistic outcome is some mess inside your own systems, visible to nobody but you, undone in a single call.
- Reading, without limit. Listing tickets, pulling one in full with its thread, searching the knowledge base, querying logs, looking up a customer's history —
list_tickets,get_ticket,search_articles,list_log_issues. An agent that read the queue at 3 AM cost you nothing and knows more than you do at 8. - Tagging and classification.
add_ticket_tagandtag_customerare the definition of cheap-to-be-wrong — a bad tag is oneremove_ticket_tagaway from never having happened. If overnight triage mislabels a few billing questions as bugs, your Tuesday sweep is slightly longer. - Internal notes.
add_ticket_notewrites staff-only text that is never emailed, never shown to the customer, never changes status. An agent leaving this is the third report of the same PDF bug — see issue 41 on a thread has done real work at zero risk. - Priority, category and assignee. Reprioritising and routing send nothing to anyone. A misrouted ticket is annoying — it is not a message.
The test for tier one is blunt. If the agent got every single one of these wrong overnight, your morning would involve a bulk correction and some swearing — and not one customer would ever know it happened. That is the whole bar. Anything that clears it should run unsupervised; supervising it is just you doing data entry.
Tier two: wrong is cheap only because you read it next
The second tier is preparation — drafted replies, drafted knowledge-base articles, a written handover of what came in overnight. The leverage is genuine and underrated: reading a draft and fixing two sentences is enormously faster than facing an empty reply box at the start of the day.
The catch is subtle. A draft is only safe while it stays a draft, and the thing that converts it is one click by a tired person. So tier two has a failure mode with nothing to do with accuracy — volume. Forty drafts in a queue you skim in four minutes is worse than zero drafts, because two of them needed you and you approved them at the same speed as the other thirty-eight. Preparation you cannot afford to read is not preparation.
So cap it. A run that drafts replies for the six tickets it is most confident about, and leaves a note saying these four are beyond me and here is why, beats one that drafts everything. The honest concession: if your queue really is forty tickets a night, drafting is not your problem — you need routing and a knowledge base before you need a night shift.
Tier three: anything that puts words in front of a customer
The third tier is short and it is absolute: anything that reaches a real person unsupervised, does not.
Ten tools in Helmdesk's MCP surface reach a customer, and the category is worth learning by its shape rather than its list. reply_to_ticket posts a staff message and emails it immediately. send_email, send_custom_email and send_email_batch put mail in inboxes. resend_email does it a second time. request_feedback and reply_to_feedback reach the same people from a different table. None can be recalled.
The reason to hold this line is not hallucination — models are good enough at support replies that fluency stopped being the risk. The reason is context the agent could not have had. It did not know you shipped a fix at 11 PM that makes the drafted answer wrong. It did not know this customer already got two apology emails this week and a third will read as sarcasm. It did not know the person asking about a refund is the one whose contract you are renegotiating. None of that is detectable from the ticket text — so no confidence score protects you from it.
And crucially: with an outbound mistake, the customer finds out before you do. That inversion is the entire tier. Every other kind of error waits politely in a database until you look. This one goes and tells someone.
The status change wearing a costume
Now the trap, and it is the single most requested overnight automation there is: just close the ones that are already answered.
It sounds like tier one. It is a status field — nothing composed, nothing written, a value changing from open to resolved. Except that resolving a ticket emails the customer a satisfaction survey. That is what resolved means: this thread is finished, here is your chance to rate it. So a tidy-up across forty answered tickets at 3 AM is forty emails at 3 AM asking people to rate support some of them will not agree they received. bulk_update_tickets makes it worse by being efficient — one call, up to a hundred surveys.
The same costume turns up in the review queue. Approving a drafted reply is what sends it: approve_agent_item reads like queue hygiene and behaves like a send button. Both are outbound actions wearing a status change's clothing, and both sit on the list of things that reach a real person for exactly that reason.
The rule that survives contact with reality is to classify by what leaves the building, not by what the verb sounds like — then enforce it somewhere the agent cannot argue with, which means the API key. A key issued without the emails:send scope makes six of those ten tools simply not exist for the session; leaving out tickets:write removes the rest. That is a boundary. A paragraph in a prompt telling the agent to be careful is not. This is also perfectly reasonable to build yourself — scoped credentials are older than agents. They just matter more now.
The bottom line
You do not need a confidence threshold. You need a consequence table. Let the agent read and label and annotate freely, let it prepare a small number of things for you to check, and let nothing reach a customer without a human in the loop who knows what shipped last night.
The question was never whether it can. It is what a mistake costs, and who finds out first — and for exactly one tier of work, the answer to the second half is not you.
Give an agent the queue, not the send button
Scoped API keys, a sandbox plane, and machine-readable annotations that tell your client which tools reach a real person.