Introducing Auto Routing on the TrueFoundry AI Gateway
Meet TrueForge: The open-source, vendor-neutral agent harness. 50% lower cost. Explore Now→
Try TrueFoundry — Live, Right Now
Get instant access to a live TrueFoundry environment. Deploy models, route LLM traffic, and explore the full platform — your sandbox is ready in seconds, no credit card required.
Introducing Auto Routing on the TrueFoundry AI Gateway
Published: September 9, 2026
Built for Speed: ~10ms Latency, Even Under Load
Blazingly fast way to build, track and deploy your models!
- Handles 350+ RPS on just 1 vCPU — no tuning needed
- Production-ready with full enterprise support
We’ve just added a new capability to the TrueFoundry AI Gateway. The new Auto-Routing feature reads each incoming request, sorts it into one of three complexity tiers, and sends it to the most appropriate model assigned to that tier.
What Auto Routing actually does
Three tiers behind a single model name
| Tier | The requests that land here | Point it at |
|---|---|---|
| simple | Quick answers, lookups, classification, simple rewrites | Your fastest, cheapest model |
| medium | Multi-step summaries, drafting, standard code, light reasoning | A balanced mid-tier model |
| complex | Deep reasoning, long context, hard problems, complex coding | Your most capable model |
How is a complexity tier selected?
There are two methods in which each query gets categorized into one of the three tiers:
Heuristic classification (the default)
- The heuristic classifier reads the prompt, scores it against a fixed set of signals, and picks a tier.
- It determines the complexity of a query based on certain parameters:
- Code
- Explicit reasoning requests
- Technical vocabulary
- Prompt length
- Multi-step structure
- Many questions at once
LLM classification
- Choose the classifier model, which must be a catalog chat model rather than another virtual model.
- Two fallback options are available: heuristic or static. A classifier failure never fails the request.
Stand-out capabilities of Auto Routing
Auto Routing goes further than simple routing rules, giving you additional capabilities to ensure model spends are optimized, while retaining accuracy:
- Escalation and fallback: Work through higher tiers if initial targets fail.
- Conversation pinning: Adjust the classification of subsequent turns based on prior responses.
- Observability: Monitor and demonstrate routing decisions and mix of tier usages to various stakeholders.
What the benchmarks show
| Workload | Cost savings | Quality retained |
|---|---|---|
| Graded academic benchmarks (11 datasets, 550 prompts) | 69% | 98% |
| Production-shaped traffic (chat, developer, agent) | 80% | Cost-only |
Getting started
Auto Routing lives on a virtual model. Go to AI Gateway → Models → Virtual Model, choose Complexity under routing type, pick a classification strategy, and set a target for each tier you want to serve. Then point your application at the virtual model's name and test it out with prompts of varying complexity in the playground.