Sign InOpen Brain
AI EngineerVideoSource Linked

Trading Desks to Clinical Trials: Parallels in Applied Vertical AI — Ayush Bhardwaj, Allos AI

Vertical agents need narrow jobs, proprietary data, observability, and expert judgment. Generic models and self-grading cannot establish whether domain-specific output is actually useful.

AI Engineer · Aug 19, 2026
Open Source Open MarkdownOpen JSON
Source Summary

Across finance and pharma, the proposed recipe starts with a **narrow task**, adds **proprietary data**, models an expert’s process, and instruments the result. A **domain expert** then judges failures; **error analysis** is presented as the cheapest first iteration method.

Practical Implication

For a specialized agent, recruit an actual user before polishing infrastructure. Use their source rankings, workflow order, and judgment to build evals and curate internal records that general models cannot access.

Agent-Ready Context
Across finance and pharma, the proposed recipe starts with a **narrow task**, adds **proprietary data**, models an expert’s process, and instruments the result. A **domain expert** then judges failures; **error analysis** is presented as the cheapest first iteration method.

For a specialized agent, recruit an actual user before polishing infrastructure. Use their source rankings, workflow order, and judgment to build evals and curate internal records that general models cannot access.

The talk argues that models cannot reliably verify their own value in domains without answer keys. Rubric-based grading can become an echo chamber, and the agent remains an assistant when causal judgment still belongs to the expert.
Connected Context · Feed7 Judgment

This narrows vertical-agent development to a user-led loop: choose one job, encode an expert’s actual workflow and proprietary evidence, instrument it, then begin with error analysis. It also draws a firm limit around model-based grading where no answer key exists: rubrics and self-verification cannot replace the expert’s causal judgment, so the system remains assistive.

Reinforcement Learning without Verifiable Rewards — Will Brown, Prime IntellectBoth address domains without deterministic answers; the simulation approach offers provisional supervision, while this Signal explains why expert judgment is still required to prevent rubric echo chambers.SimulationMaxxing: How we ship agents 20× faster — Aman Gupta (Nubank) + Shreya Rajpal (Snowglobe)Simulated traces can expand pre-production evaluation for the narrow workflow, but their required expert and production validation matches this Signal’s warning that models cannot establish domain value themselves.From Agent Traces to Agent Simulations — Rustem Feyzkhanov, Snorkel AIReplayable production-derived environments provide the instrumentation and fixed comparisons this recipe needs, while expert outcome judgment remains necessary where task success has no definitive answer key.200 Million Patient Interactions Later — Vivek Muppalla, Hippocratic AIHippocratic AI’s specialist checks and offline verification show an implementation consequence of modeling narrow expert processes: reliability should be distributed across purpose-built controls rather than entrusted to one agent.
Context Map
agentdataresearch#harness-engineering#agent-evals#agent-reliability
Uncertainty
The talk argues that models cannot reliably verify their own value in domains without answer keys. Rubric-based grading can become an echo chamber, and the agent remains an assistant when causal judgment still belongs to the expert.