Meet TrueForge: The open-source, vendor-neutral agent harness. 50% lower cost. [Explore Now→](https://github.com/truefoundry/trueforge)

Try TrueFoundry — Live, Right Now

Get instant access to a live TrueFoundry environment. Deploy models, route LLM traffic, and explore the full platform — your sandbox is ready in seconds, no credit card required.

# Introducing Auto Routing on the TrueFoundry AI Gateway

[By Rhea Jain](/content/blogs/authors/rhea-jain/index.html)

Published: September 9, 2026

### Built for Speed: ~10ms Latency, Even Under Load

Blazingly fast way to build, track and deploy your models!

- Handles 350+ RPS on just 1 vCPU — no tuning needed
- Production-ready with full enterprise support

We’ve just added a new capability to the TrueFoundry AI Gateway. The new Auto-Routing feature reads each incoming request, sorts it into one of three complexity tiers, and sends it to the most appropriate model assigned to that tier.

## What Auto Routing actually does

### Three tiers behind a single model name

| Tier   | The requests that land here                            | Point it at                        |
|--------|--------------------------------------------------------|------------------------------------|
| simple | Quick answers, lookups, classification, simple rewrites | Your fastest, cheapest model       |
| medium | Multi-step summaries, drafting, standard code, light reasoning | A balanced mid-tier model          |
| complex | Deep reasoning, long context, hard problems, complex coding | Your most capable model            |

## How is a complexity tier selected?

There are two methods in which each query gets categorized into one of the three tiers:

1. **Heuristic classification (the default)**
   - The heuristic classifier reads the prompt, scores it against a fixed set of signals, and picks a tier.
   - It determines the complexity of a query based on certain parameters:
     - **Code**
     - **Explicit reasoning requests**
     - **Technical vocabulary**
     - **Prompt length**
     - **Multi-step structure**
     - **Many questions at once**

2. **LLM classification**
   - Choose the classifier model, which must be a catalog chat model rather than another virtual model.
   - Two fallback options are available: heuristic or static. A classifier failure never fails the request.

## Stand-out capabilities of Auto Routing

Auto Routing goes further than simple routing rules, giving you additional capabilities to ensure model spends are optimized, while retaining accuracy:

- **Escalation and fallback**: Work through higher tiers if initial targets fail.
- **Conversation pinning**: Adjust the classification of subsequent turns based on prior responses.
- **Observability**: Monitor and demonstrate routing decisions and mix of tier usages to various stakeholders.

## What the benchmarks show

| Workload                                               | Cost savings | Quality retained |
|--------------------------------------------------------|--------------|------------------|
| Graded academic benchmarks (11 datasets, 550 prompts) | 69%          | 98%              |
| Production-shaped traffic (chat, developer, agent)     | 80%          | Cost-only        |

## Getting started

Auto Routing lives on a virtual model. Go to **AI Gateway** → **Models** → **Virtual Model**, choose **Complexity** under routing type, pick a classification strategy, and set a target for each tier you want to serve. Then point your application at the virtual model's name and test it out with prompts of varying complexity in the playground.
