Pick models with proof, not vibes
Benchmark candidate models and prompts against each workflow's real tests, scored by an LLM judge. Get a routing recommendation and catch drift before it ships.
Scored on your prompts, not leaderboards
Cran runs every prompt x model pair, scores outputs with class-aware judge rubrics, and recommends a primary, fallbacks, and a cost route you can publish live.
| Model | Score | P95 | Cost | Route |
|---|---|---|---|---|
| openai/gpt-5.4-mini | 4.6 | 640ms | $0.004 | primary |
| anthropic/claude-sonnet-4-6 | 4.5 | 920ms | $0.012 | fallback |
| google/gemini-2.5-flash | 4.3 | 580ms | $0.003 | fallback |
| openai/gpt-5.5 | 4.7 | 1.4s | $0.048 | skipped |
Harness optimization
Rewrites a workflow's prompt leaner, re-tests it, and only recommends the swap when tokens drop and quality holds within a non-inferiority bound.
Architecture review
Reads all workflows and flags prompt bloat, unused tools, caching opportunities, and more - ten finding types across four severities.
Drift detection
Re-checks on a schedule and flags when a model regresses on your tests.
One-click publish
Turn a recommendation into live routing from the Optimize inbox.
Explore the rest of the control plane
Map your AI in 2 minutes.
Connect your repo, route every call through Cran, and publish routing when you have proof - not vibes.