Avoid choosing a winner first
Performance depends on the task, context and instructions. Run a blinded comparison using your own price objections, incomplete requests and handoff scenarios. Hide model names from reviewers to reduce brand preference.
Test real Arabic usage
Evaluate Egyptian phrasing, names, numbers and mixed-language messages. Include imperfect voice transcripts. Reward necessary clarification: a polite reply with the wrong price is worse than a concise request for missing information.
Keep the decision revisable
Compare latency, resolved-task cost and tool accuracy alongside language quality. Different tasks may justify different models, but routing complexity needs a measured benefit. Reuse the evaluation set after prompt, model or knowledge changes.