Direct answer
Choose models per task using a representative evaluation set and an operating threshold for quality, latency, cost, privacy, and failure behavior. Route to the least expensive model that reliably meets the requirement, then escalate only when the input or consequence justifies it. Model routing is a policy backed by evidence—not a permanent hierarchy of “smart” and “cheap.”
One model is simple until the workload diversifies
An AI product may classify support tickets, extract fields, retrieve documents, draft messages, analyze complex cases, and review high-risk actions. Those tasks do not have the same difficulty or consequence.
Sending everything to the largest model raises cost and latency. Sending everything to the smallest one creates rework and hidden quality loss. A router should reflect the job being done.
Define task classes before model classes
Group requests by acceptance criteria, not by interface. “Chat” is not one task. It may include factual retrieval, policy explanation, free-form brainstorming, and an instruction to change account state.
For each task class, define:
- required quality dimensions and minimum scores;
- maximum latency and cost per successful outcome;
- context length and modality needs;
- data-processing or residency constraints;
- allowed tools and autonomy;
- fallback and escalation behavior.
Only then compare candidate models.
Evaluate the whole route
Run the same representative cases through each model and surrounding prompt, retrieval, tools, and validation. Measure task completion, groundedness, format compliance, tool accuracy, latency, usage, and human correction.
Cost per call is not the useful unit. Measure cost per accepted outcome. A cheaper model that causes more retries, escalations, or operator edits may cost more in practice.
Inspect tail behavior. A model with a strong average but catastrophic failures on one high-value class may be inappropriate for that route.
Use escalation deliberately
A common pattern is a fast first pass followed by escalation when deterministic checks fail, confidence is low, evidence conflicts, or the action carries higher consequence. The larger model receives the hard cases rather than every case.
Escalation should not become an invisible retry loop. Set a maximum number of routes, preserve the reason for escalation, and stop for human review when the evidence remains insufficient.
The router can also select non-model behavior. A cached answer, rules engine, template, database query, or refusal may be the right route.
Version routing policy like code
Provider behavior, pricing, context limits, and availability change. Keep routing rules in configuration or code, attach versions to traces, and run evals before changing them.
Monitor quality and outcome metrics by route after deployment. Input distributions drift. A class that once contained short support messages may begin receiving long attachments, new languages, or adversarial content.
Fallbacks need testing too. If the primary provider fails, the secondary model may use a different schema, tool format, safety behavior, or tokenization pattern. “Switch providers” is not a recovery plan unless the product has validated the switch.
Architecture creates the savings
Zenveus used layered cost controls in a regulatory AI platform: change detection stopped unchanged sources before paid processing, batch extraction handled bulk work, prompt caching reduced repeated context cost, and local embeddings removed a separate paid embedding call. The system reserved expensive generation for the part that required it.
That is model routing in the broader sense. The best request is sometimes the request the architecture avoided.
The acceptance test
For every route, the team should be able to show the evaluation set, required threshold, observed quality, p50 and p95 latency, cost per accepted outcome, escalation rate, and known failure modes.
If a route exists because one model “felt good in testing,” it is an opinion wired into production.
Related Zenveus services: Agentic AI Development and Elastic Infrastructure