Don’t Use Their Models.
Use Yours.Don’t Use Their Models.Use Yours.
Compare model capabilities head-to-head, then deploy the most efficient.
Understudy Optimizes the Complete LLM Production Route
Test. Learn. Deploy. Repeat.
Capture
Capture traces from LLM production workflows with a single install that deploys within coding agents you already use. Hosted infrastructure is optional.
Evaluate
Evaluate the captured traces and set a benchmark for success. Every future model switch meets or exceeds it in A/B testing.
Train
Train and fine-tune a new model on prompts and weights you always own. Start locally and scale into cloud-hosted processes as you see success.
Deploy
Deploy a new model only when the held-out eval is beaten. Serve it wherever you want. Production data feeds back into training and compounds performance over time.
Proof of performance
1.13x Sonnet score at 25% of Sonnet cost.
Read more →Qwen3-8B tuned to match Sonnet performance with almost equivalent reliability at 6.0x lower measured token cost.
Read more →Sentiment analysis: an Understudy post-trained 30B open model labeled 39,962 comments at 4.4x lower full-table cost than Sonnet and 50x lower cost than Opus.
Read more →Frequently asked questions
What does Understudy optimize?
Understudy optimizes complete production routes for repeated LLM work: the harness, model, and supply path. That includes prompts, schemas, tool-call adapters, reasoning mode, token caps, scorers, retry policy, batching, context compaction, parsers, model choice, fine-tuned descendants, and serving path.
Do we need to move our workflow into a hosted app?
No. The CLI, MCP server, skills, and local workbench run inside the coding agents and environments your team already uses. Hosted infrastructure is optional when an optimization needs cloud training or serving.
How does Understudy know whether a smaller model is good enough?
The system turns production traces and expert review into evals. A cheaper route only replaces a frontier baseline after it satisfies the task-specific quality bar, with failures and uncertainty escalated, repaired, or converted into training data.
When should we keep using a frontier model?
Keep the frontier model where premium capability changes the outcome. For routine agentic operations, the goal is enough intelligence at the right latency and price: classify a message, choose a tool, fill structured arguments, repair malformed calls, or create signal for later optimization.
Do we own the resulting models?
Yes. The goal is to hand off prompts, evaluators, routing rules, and specialist model weights that your team can serve on Fireworks, Bedrock, Vertex, or your own GPUs.
What teams are a fit for private preview?
The best fit is a team with a real production LLM workload, meaningful cost or latency pressure, repeated task volume, and domain experts who can review outputs.
Interested?
Understudy is in private preview with a small group of design partners. We work closely with each team to install the proxy, capture traces, and train the first replacement model alongside their domain experts.
We are looking for production LLM workflows where cost or latency is starting to hurt and the people who know what good looks like are not on an ML team.