Trading Desks to Clinical Trials: Parallels in Applied Vertical AI — Ayush Bhardwaj, Allos AI
Vertical agents need narrow jobs, proprietary data, observability, and expert judgment. Generic models and self-grading cannot establish whether domain-specific output is actually useful.
Across finance and pharma, the proposed recipe starts with a **narrow task**, adds **proprietary data**, models an expert’s process, and instruments the result. A **domain expert** then judges failures; **error analysis** is presented as the cheapest first iteration method.
For a specialized agent, recruit an actual user before polishing infrastructure. Use their source rankings, workflow order, and judgment to build evals and curate internal records that general models cannot access.
Across finance and pharma, the proposed recipe starts with a **narrow task**, adds **proprietary data**, models an expert’s process, and instruments the result. A **domain expert** then judges failures; **error analysis** is presented as the cheapest first iteration method. For a specialized agent, recruit an actual user before polishing infrastructure. Use their source rankings, workflow order, and judgment to build evals and curate internal records that general models cannot access. The talk argues that models cannot reliably verify their own value in domains without answer keys. Rubric-based grading can become an echo chamber, and the agent remains an assistant when causal judgment still belongs to the expert.
This narrows vertical-agent development to a user-led loop: choose one job, encode an expert’s actual workflow and proprietary evidence, instrument it, then begin with error analysis. It also draws a firm limit around model-based grading where no answer key exists: rubrics and self-verification cannot replace the expert’s causal judgment, so the system remains assistive.