Constrain the decision

Define categories such as pricing, order tracking, complaint and unclear. Give boundary examples for similar cases. This is a useful smaller-model evaluation task without assuming equivalent performance in open-ended conversation.

Handle ambiguous cases

A message can combine price objections and a delayed order. Allow multiple labels or review instead of forcing one wrong answer. Validate against human labels rather than trusting a confidence number generated by the model.

Measure category-level quality

Track accuracy per category, review rates and costly misroutes. Missing an urgent complaint differs from escalating a general query. Test expected volume and refresh examples when services or customer vocabulary change.