Production applications rarely use just one model. As teams add providers and models, model selection, fallbacks, spending limits, and provider-specific APIs become application logic.
We’re introducing Chalk Model Gateway, an OpenAI-compatible model gateway that runs inside your Chalk deployment, in your own cloud.
Chalk Model Gateway does what you’d expect from a gateway: routing, fallback, budgets, rate limits, and tracing. But, managed gateways put another company’s infrastructure between your application and your models. Self-hosted gateways give your team another service to operate. Chalk Model Gateway gives you the benefits of both without either tradeoff: it runs inside your existing Chalk deployment, in your cloud.
With Chalk Model Gateway:
- No third-party gateway sits between your agents and the models they call.
- Your provider keys stay in your environment.
- If you serve open-weight models on Chalk, those requests don’t leave your cloud at all.
For teams with sensitive data control requirements, Chalk Model Gateway, which is deployed in your environment, can simplify security and compliance reviews.
Production agents need more than a model endpoint
A few years ago, most teams building agents were validating the idea with a single model.
That’s changing. An a16z survey of 100 CIOs at Global 2000 companies found that 81% now use three or more model families, up from 68% less than a year earlier. As agents move into production, model choice becomes a systems problem. Wiring every provider straight into application code breaks down fast.
You need to answer questions like:
- What happens when a provider rolls back a model or a model is unavailable?
- Which requests actually need the most expensive model?
- How much can each team, agent, or workflow spend?
- How do you switch providers without rewriting application code?
- Which model handled a particular request?
- What happens when you introduce a different model and quality changes?
A gateway answers these in one place.
Meet Chalk Model Gateway
Your application talks to one endpoint. The gateway handles where the request goes. You bring your own keys for the providers you already use: OpenAI, Anthropic, Jev, Google Vertex, AWS Bedrock, and more. You can also route to the endpoints for any open-weight models you serve yourself with vLLM. From there, you can put policy around model traffic:
- Fallbacks. Define an ordered list of backup models. If the primary model is unavailable or returns an error, the gateway tries the next one before your agent sees a failure.
- Budgets and rate limits. Control token usage and request volume across API keys, providers, models, and shared usage pools. Set daily, weekly, or monthly budgets alongside tokens-per-minute, requests-per-minute, and concurrency limits.
- Judge-based routing. Route by task type with a policy you configure: a local heuristic difficulty classifier, or a remote judge backed by any LLM or System One model. Start any new policy in shadow mode to see which requests it would downgrade before it touches production traffic.
- Open-model routing. Serve open-weight models on GPUs in your own cloud with Chalk Compute and route to them through the same gateway as your frontier providers.
- Tracing. Every request gets a trace, with controls over sampling and content retention. Errors are always exported.
For example, you can easily point your existing OpenAI SDK at your Chalk host:
from openai import OpenAI
client = OpenAI(
base_url="https://<your-chalk-api-host>/v1/router",
api_key="<YOUR_ROUTER_API_KEY>",
default_headers={"X-Chalk-Env-Id": "<env-id>"},
)
resp = client.chat.completions.create(
model="openai/gpt-6-luna",
messages=[{"role": "user", "content": "Hello"}],
)
print(resp.choices[0].message.content)Check out the Model Gateway docs for the full reference.
Don’t send every request to your most expensive model
Routing lets you pay for the capability each request needs. For years, the easiest strategy was to pick the strongest model and send everything to it. That made sense when model quality was changing quickly and volume was relatively low. At production scale, the economics are different.
Here is what a typical agent workload looks like at list price, blended at three input tokens for every output token:
Model
Input ($/M tokens)
Output ($/M tokens)
Blended, 3:1 ($/M tokens)
- Frontier hosted model
- 4.00
- 20.00
- 8.00
- Smaller hosted model
- 2.00
- 10.00
- 4.00
- Lower-cost, efficient hosted model
- 0.30
- 1.20
- 0.53
- Self-hosted, open-weight, fine-tuned model
- -
- -
- 10-15% of frontier hosted model costs [1]
[1] An enterprise social-dating application was running a moderation workload on OpenAI models. Their analysis found that switching to a fine-tuned, open-weight model hosted on Chalk would reduce inference costs to as little as 10-15% of their current spend, for an estimated $150k+ in annual savings.
At one billion tokens per month, sending every request to a frontier model would cost roughly $8,000 at these list prices. Now imagine a routing policy that sends 70% of tokens to a lower-cost, efficient hosted model and escalates the remaining 30% to the frontier model. The cost falls to roughly $2,770, a 65% reduction. The gateway turns model choice into a policy you can change, instead of logic hard-coded into your application.
The economics of open models can be compelling. But cheaper inference isn't enough. You need to know which requests can safely use a smaller model.
With Chalk Model Gateway, start any new routing policy in shadow mode to see which requests it would downgrade before you commit production traffic to it. You can measure the tradeoff by routing to a different model, evaluating and observing, adjusting, and rerouting again.
One stack for production agents, in your cloud
Chalk Model Gateway is one part of Chalk’s platform for production AI. The pieces work together:
- Your data and cloud. Chalk runs in your environment and provides the live context and features your agents and models need. The router is in your environment, so you can configure your application to ensure data doesn’t leave your cloud.
- Your models. Use frontier APIs with your own provider keys, or run open-weight models on your own GPUs with Chalk Compute.
- Your routing policies. Use Chalk Model Gateway to decide which model handles each request based on cost, capability, availability, and policy.
These systems increasingly depend on each other. All three operate within the same Chalk environment, so you can move from context to inference to routing without stitching together another set of services. The right model depends on the request, and the right decision depends on the context available to that model.
Chalk’s Model Gateway sits alongside the rest of Chalk’s production AI infrastructure: live context and features, model serving, compute, evaluation, and now routing. The model sets cost, the context sets quality, and the routing policy decides which model sees which request.
If you are building agentic workflows and want reliable, cost-aware model access without sending traffic outside your cloud, talk with the Chalk team today and check out the Model Gateway docs.







