One workload. Every model. Measured by your evals.

Your models get better every week.

Understudy connects one production workload to every viable model route—and keeps improving prompts, routing, and specialist weights against the evals your team trusts.

understudy.run
01 / capture
production traces indexed
02 / evaluate
Sonnet 4.6 baseline0.557Understudy route0.630
03 / promote
quality +13%·cost 0.25×
candidate ready for held-out review

The optimization thesis

Choosing a model is a moment. Improving one is a process.

01

The landscape keeps moving

New models and serving paths change the price-performance frontier every week. A hard-coded provider leaves yesterday’s decision in production.

02

Your eval defines good

Understudy sits between a real workload and every viable route. It improves prompts, compares models, and post-trains only against outcomes your team accepts.

03

The integration stays put

Keep one workload contract while the prompt, route, and specialist weights improve behind it—without rewriting the product every time the frontier shifts.

Eval score · your workload

Hill climb on your evals.

baselinetuneswappost-train
A rising evaluation score as a workload moves from its baseline through prompt tuning, model comparison, and specialist post-training.
01 / Baseline

Evals define good. Start with the frontier route your team already trusts and make its quality bar visible.

02 / Tune

Prompts improve first. Test the harness, schema, reasoning mode, and tool contract before changing the model.

03 / Swap

Models compete on the same workload. A cheaper or faster route only moves forward when it clears the held-out bar.

04 / Post-train

When the use case earns it, turn the best traces and corrections into specialist weights your team owns.

01 / pressure

Bring your eval suite—or build one from production traces. Every candidate is scored against the outcome your workload actually needs. The curve is that score; every change exists to push it up.

Price / performance

Swap models. Keep the workload.

One integration keeps working while Understudy finds a route that is smarter, faster, or cheaper for the job in front of it.

01Router

Understudy route

Give one production workload access to the route that best balances measured quality, cost, and latency.

02Evals

Quality gate

Keep the frontier baseline until a candidate clears the task-specific quality bar on held-out work.

03Measured

Cost floor

Let smaller and open models compete on the work you actually run—not a generic leaderboard.

04P95

Latency budget

Compare response time beside quality so a smarter route never quietly makes the product unusable.

05Rollback

Safe promotion

Review the evidence, ramp the winning route, and keep the prior baseline available when reality disagrees.

06Portable

No lock-in

Keep the prompts, evaluators, routing rules, and specialist model weights your workload produces.

Price / performance

Same evals. A smaller invoice.

The frontier baseline stays visible. Every candidate earns its place on the same workload, using the same eval, with measured cost and latency beside it.

Make agents smarter
+13% higher eval score
vs Sonnet 4.6
Sonnet 4.60.557
Baseline open model0.400
Understudy ladder0.630

1.13x Sonnet score at 25% of Sonnet cost.

Read more →
Make product features faster
5.2x lower latency
Sonnet 4.61.935s
Understudy route (8B)369ms

Qwen3-8B tuned to match Sonnet performance with almost equivalent reliability at 6.0x lower measured token cost.

Read more →
Make product features economically viable
50x lower cost
Understudy ladder$2.82
Sonnet$12.48
Opus$139.63

Sentiment analysis: an Understudy post-trained 30B open model labeled 39,962 comments at 4.4x lower full-table cost than Sonnet and 50x lower cost than Opus.

Read more →
See all product use-cases →

Product

Inside the platform.

Quality, latency, and cost stay on the same screen so every route change has evidence—and every result can be challenged.

Eval dashboardCRM agent · measured run
Candidate score
0.630
+13% vs frontier baseline
capturetunepromote
Model routing
Sonnet 4.60.557
Understudy route0.630
Cost analytics
Baseline
$1.12
Candidate
$0.27

Same documented CRM workload, measured beside its frontier baseline.

+13%
higher eval score
CRM workload
5.2×
lower latency
operations workload
50×
lower cost
39,962 sentiment labels

The details

A clearer path to your model.

What does Understudy optimize?

Understudy optimizes complete production routes for repeated LLM work: the harness, model, and supply path. That includes prompts, schemas, tool-call adapters, reasoning mode, token caps, scorers, retry policy, batching, context compaction, parsers, model choice, fine-tuned descendants, and serving path.

Do we need to move our workflow into a hosted app?

No. The CLI, MCP server, skills, and local workbench run inside the coding agents and environments your team already uses. Hosted infrastructure is optional when an optimization needs cloud training or serving.

How does Understudy know whether a smaller model is good enough?

The system turns production traces and expert review into evals. A cheaper route only replaces a frontier baseline after it satisfies the task-specific quality bar, with failures and uncertainty escalated, repaired, or converted into training data.

When should we keep using a frontier model?

Keep the frontier model where premium capability changes the outcome. For routine agentic operations, the goal is enough intelligence at the right latency and price: classify a message, choose a tool, fill structured arguments, repair malformed calls, or create signal for later optimization.

Do we own the resulting models?

Yes. The goal is to hand off prompts, evaluators, routing rules, and specialist model weights that your team can serve on Fireworks, Bedrock, Vertex, or your own GPUs.

What teams are a fit for private preview?

The best fit is a team with a real production LLM workload, meaningful cost or latency pressure, repeated task volume, and domain experts who can review outputs.

Private preview

Start with one workload. Improve forever.

Bring a production LLM workflow where cost or latency is starting to hurt. We will help your team measure it and find the first specialist route worth owning.