Evals

Pick models with proof, not vibes

Benchmark candidate models and prompts against each workflow's real tests, scored by an LLM judge. Get a routing recommendation and catch drift before it ships.

Scored on your prompts, not leaderboards

Cran runs every prompt x model pair, scores outputs with class-aware judge rubrics, and recommends a primary, fallbacks, and a cost route you can publish live.

Six class-aware judge rubrics Quality-vs-cost routing recommendations Drift alerts on regressions
Audits & evals
refund_classifier12× cost spread at similar quality
ModelScoreP95CostRoute
openai/gpt-5.4-mini4.6640ms$0.004primary
anthropic/claude-sonnet-4-64.5920ms$0.012fallback
google/gemini-2.5-flash4.3580ms$0.003fallback
openai/gpt-5.54.71.4s$0.048skipped

Harness optimization

Rewrites a workflow's prompt leaner, re-tests it, and only recommends the swap when tokens drop and quality holds within a non-inferiority bound.

Architecture review

Reads all workflows and flags prompt bloat, unused tools, caching opportunities, and more - ten finding types across four severities.

Drift detection

Re-checks on a schedule and flags when a model regresses on your tests.

One-click publish

Turn a recommendation into live routing from the Optimize inbox.

Map your AI in 2 minutes.

Connect your repo, route every call through Cran, and publish routing when you have proof - not vibes.