Avoid choosing a winner first

Performance depends on the task, context and instructions. Run a blinded comparison using your own price objections, incomplete requests and handoff scenarios. Hide model names from reviewers to reduce brand preference.

Test real Arabic usage

Evaluate Egyptian phrasing, names, numbers and mixed-language messages. Include imperfect voice transcripts. Reward necessary clarification: a polite reply with the wrong price is worse than a concise request for missing information.

Keep the decision revisable

Compare latency, resolved-task cost and tool accuracy alongside language quality. Different tasks may justify different models, but routing complexity needs a measured benefit. Reuse the evaluation set after prompt, model or knowledge changes.