The landscape keeps moving
New models and serving paths change the price-performance frontier every week. A hard-coded provider leaves yesterday’s decision in production.
One workload. Every model. Measured by your evals.
Understudy connects one production workload to every viable model route—and keeps improving prompts, routing, and specialist weights against the evals your team trusts.
The optimization thesis
New models and serving paths change the price-performance frontier every week. A hard-coded provider leaves yesterday’s decision in production.
Understudy sits between a real workload and every viable route. It improves prompts, compares models, and post-trains only against outcomes your team accepts.
Keep one workload contract while the prompt, route, and specialist weights improve behind it—without rewriting the product every time the frontier shifts.
Eval score · your workload
Evals define good. Start with the frontier route your team already trusts and make its quality bar visible.
Prompts improve first. Test the harness, schema, reasoning mode, and tool contract before changing the model.
Models compete on the same workload. A cheaper or faster route only moves forward when it clears the held-out bar.
When the use case earns it, turn the best traces and corrections into specialist weights your team owns.
Bring your eval suite—or build one from production traces. Every candidate is scored against the outcome your workload actually needs. The curve is that score; every change exists to push it up.
Price / performance
One integration keeps working while Understudy finds a route that is smarter, faster, or cheaper for the job in front of it.
Give one production workload access to the route that best balances measured quality, cost, and latency.
Keep the frontier baseline until a candidate clears the task-specific quality bar on held-out work.
Let smaller and open models compete on the work you actually run—not a generic leaderboard.
Compare response time beside quality so a smarter route never quietly makes the product unusable.
Review the evidence, ramp the winning route, and keep the prior baseline available when reality disagrees.
Keep the prompts, evaluators, routing rules, and specialist model weights your workload produces.
Price / performance
The frontier baseline stays visible. Every candidate earns its place on the same workload, using the same eval, with measured cost and latency beside it.
1.13x Sonnet score at 25% of Sonnet cost.
Read more →Qwen3-8B tuned to match Sonnet performance with almost equivalent reliability at 6.0x lower measured token cost.
Read more →Sentiment analysis: an Understudy post-trained 30B open model labeled 39,962 comments at 4.4x lower full-table cost than Sonnet and 50x lower cost than Opus.
Read more →Product
Quality, latency, and cost stay on the same screen so every route change has evidence—and every result can be challenged.
Same documented CRM workload, measured beside its frontier baseline.
The details
Understudy optimizes complete production routes for repeated LLM work: the harness, model, and supply path. That includes prompts, schemas, tool-call adapters, reasoning mode, token caps, scorers, retry policy, batching, context compaction, parsers, model choice, fine-tuned descendants, and serving path.
No. The CLI, MCP server, skills, and local workbench run inside the coding agents and environments your team already uses. Hosted infrastructure is optional when an optimization needs cloud training or serving.
The system turns production traces and expert review into evals. A cheaper route only replaces a frontier baseline after it satisfies the task-specific quality bar, with failures and uncertainty escalated, repaired, or converted into training data.
Keep the frontier model where premium capability changes the outcome. For routine agentic operations, the goal is enough intelligence at the right latency and price: classify a message, choose a tool, fill structured arguments, repair malformed calls, or create signal for later optimization.
Yes. The goal is to hand off prompts, evaluators, routing rules, and specialist model weights that your team can serve on Fireworks, Bedrock, Vertex, or your own GPUs.
The best fit is a team with a real production LLM workload, meaningful cost or latency pressure, repeated task volume, and domain experts who can review outputs.
Private preview
Bring a production LLM workflow where cost or latency is starting to hurt. We will help your team measure it and find the first specialist route worth owning.