How to Automate Support Triage With an LLM, Safely
An LLM is great at sorting support tickets and dangerous at answering them unsupervised. Here is the safe split.
Support teams drown in tickets, and an LLM looks like the rescue. The trap is that the same model does two very different jobs at once: sorting a ticket and answering it. Sorting is low-risk and easy to check. Answering, unsupervised, is how you refund the wrong customer or promise something your product does not do. The safe move is to split those jobs and treat them as separate systems with separate rules.
#Triage and resolution carry different risk
Triage is reading and labeling. What is this ticket about, how urgent is it, which queue does it belong in. If the model gets a label wrong, a human in that queue catches it and moves it. The blast radius is one misrouted ticket, and it is cheap to undo.
Resolution is acting on the customer's behalf. Sending a reply, issuing a refund, resetting an account. A wrong action here reaches the customer directly and is often hard to reverse. Same model, completely different consequences when it is wrong. Once you see them as separate, the design writes itself: automate the reading, gate the acting.
#What the model should actually output
Do not let triage return free text. Make it return a small, closed set of structured fields the rest of your system can trust. Category from a fixed list. Priority from P1 to P4. A boolean for needs_human. A confidence score.
Closed outputs are the whole point. If the category must be one of twelve values, a hallucinated thirteenth fails a schema check instead of silently routing a billing issue to the returns team. You can validate "refund_request". You cannot validate a paragraph. It also makes the system testable. Build a set of 200 real tickets with known correct labels and measure accuracy on every prompt change, which you can never do cleanly with freeform answers.
#Draft, do not send
There is a useful middle ground between pure triage and full automation, and it is drafting. The model writes a suggested reply and a human approves, edits, or rejects it before anything goes out. The agent stays fast and the customer never sees an unreviewed answer.
The honest tradeoff: this does not remove the human, it speeds them up. If someone sells you "fully autonomous support resolution," they are either handling only trivial tickets or shipping wrong answers they have not measured yet. The one place to allow true auto-send is a narrow, well-tested band: high confidence, a known-safe category like a password reset link, and an action that is cheap to reverse. Everything else waits for a person.
#Route by confidence, and mean it
Confidence only helps if low confidence changes what happens. A score the system ignores is decoration. Wire it to routing. Above your threshold, the ticket flows on the fast path. Below it, the ticket goes straight to a human with no automated action attempted.
Calibrate the threshold against real tickets. Run a batch, find where the model was confident and wrong, and set the line above that. Then watch the escalation rate. If almost nothing escalates, your threshold is too loose and errors are leaking through. A triage system that sends its uncertain cases to a person is working correctly. That escalation path is the safety valve, so build it first, before you tune anything else.
#The bottom line
Automating support triage is a real win, and it is safe because sorting is easy to check and easy to undo. The danger starts the moment you let the same model answer without a human between it and the customer. Draw that line clearly, route uncertainty to a person, and only auto-send in a narrow band you have actually measured. Do that and you get the throughput without the incident report.