Constrain the decision
Define categories such as pricing, order tracking, complaint and unclear. Give boundary examples for similar cases. This is a useful smaller-model evaluation task without assuming equivalent performance in open-ended conversation.
Handle ambiguous cases
A message can combine price objections and a delayed order. Allow multiple labels or review instead of forcing one wrong answer. Validate against human labels rather than trusting a confidence number generated by the model.
Measure category-level quality
Track accuracy per category, review rates and costly misroutes. Missing an urgent complaint differs from escalating a general query. Test expected volume and refresh examples when services or customer vocabulary change.