Preferences Over Benchmarks: Model Routing — Archana Kamath & Tyler Gillam, DigitalOcean
Per-task model routing cut the demonstrated coding session’s cost from 44¢ to 14¢ with similar completion time, but builders still need workload-specific evals to validate quality.
DigitalOcean’s demo routed code generation, testing, and documentation to configured models while a comparison session used Opus. The routed session cost **14 cents versus 44 cents**, and its routing decision was reported at **under 200 milliseconds**.
Define routing around task, latency, cost, preferred models, and hard failover rules. Evaluate the full coding workflow on your own cases, then adjust model pools and policies instead of choosing one model from a public leaderboard.
DigitalOcean’s demo routed code generation, testing, and documentation to configured models while a comparison session used Opus. The routed session cost **14 cents versus 44 cents**, and its routing decision was reported at **under 200 milliseconds**. Define routing around task, latency, cost, preferred models, and hard failover rules. Evaluate the full coding workflow on your own cases, then adjust model pools and policies instead of choosing one model from a public leaderboard. The app comparison was a live demonstration and visual judgment, not a controlled quality study. A separate evaluation scored router correctness at 90% versus 95% for Opus, which the speakers described as close, but no judge uncertainty was quantified.
This adds a workflow-level demonstration that routing different coding stages can reduce observed cost and keep routing latency small, while confirming that model choice should follow local preferences and cases rather than leaderboards. It does not establish equivalent output quality: the comparison was visual, and the reported router scores lack judge-uncertainty analysis, so the result supports experimentation and policy tuning rather than a default route.