A single customer message can contain a technical question, a complaint and a request for a discount. Passing it to a general-purpose agent without distinction leaves important decisions implicit: which information to consult, which team to involve and which commitments are authorised?

A classifier assigns one or more categories to a request. Routing uses that result to select the next workflow. Classifying a message as a complaint does not, by itself, authorise a refund.

Start with business responsibilities

Imagine a fictional support department. We could define four treatments: a documentation response, a technical incident, a commercial request and review of a sensitive case. An ambiguous request should be able to enter a review queue instead of being forced into an unsuitable category.

Before choosing a model, define categories and responsibilities. What happens when several topics coexist? Who takes over a misrouted request? Which signals require review? These are operating decisions for the service.

Choose the complexity the task needs

An explicit rule may be sufficient for some cases; a classification model or LLM may help when wording varies. The important step is evaluating results on the requests actually received. We recommend comparing options against a simple baseline before adding complexity.

Teams should label a set of examples and discuss disagreements. Keep separate examples for evaluation, include rare but important requests, and avoid testing only on wording already used to configure the system. Record why a case belongs in a category so the review remains reproducible.

A score is not enough to decide

A confidence score is useful only when its behaviour has been checked. A calibrated probability should match observed frequencies on comparable cases. A confidence number generated spontaneously by an LLM is not automatically a calibrated probability.

A routing threshold should reflect the cost of mistakes and the acceptable review workload. We advise against choosing a threshold simply because it “looks high”. Measure what passes, what gets blocked and which categories perform poorly. Sensitive cases may require review regardless of the score.

Make escalation useful

An escalation should include the request, retrieved evidence, proposed route and reason for uncertainty. It needs an assigned recipient and a defined handling time. Otherwise, automation merely moves work into an invisible queue.

In this example, a documentation request could lead to an evidence-backed draft. An issue involving equipment safety would follow the incident process. A commercial discount would remain subject to the company’s decision-making authority. Each branch should have clear permissions as well as a clear owner.

Deliverables for a serious project

Mintera proposes to map requests and decisions, assemble a reference set, test routing and design handoffs to teams. Measurement should cover each category: correct routing, serious mistakes, review rate, time to resolution and total handling cost.

After deployment, new products and new requests may change the situation. An owner should review errors, update examples and retest settings. Comparing reviewed cases with automatically handled ones can also reveal gaps that a headline accuracy figure would hide.

The classifier becomes a monitored component of the process, with an explicit decision scope and a practical way for staff to correct it.

Map your teams’ requests and decisions.