Regulated decisions, under audit
Document AI in mortgage underwriting, where an extraction confidence is not a legally sufficient reason. Adverse-action notices, exception queues, and an immutable trail a regulator will eventually read.
I work forward-deployed, which usually means being the one person between an executive's whiteboard and the code that ends up under audit, in regulated industries like pharma, healthcare, and lending. Nine years in, and the model is rarely the hard part. What decides whether an agent survives production is the stuff nobody puts in a demo: the eval that catches a regression before a user does, the tracing that turns a 3am failure into something you can actually read, whether the client's team can still run it after I've gone.
That last part is the one I care about most. And the week right after launch is still my favorite of any project, when real traffic starts arguing with assumptions you didn't know you'd made. This is where I write about it.

Document AI in mortgage underwriting, where an extraction confidence is not a legally sufficient reason. Adverse-action notices, exception queues, and an immutable trail a regulator will eventually read.
Agent fleets against Workday, Salesforce, ServiceNow, SAP and Palantir. Their permission models grant to people holding jobs, and an agent is neither, so least privilege has no native expression.
Google, Snowflake and Databricks agent platforms federated over A2A and MCP. Each is sound alone; the failures live in the joins, where three catalogs disagree and the loosest one decides.
A division-wide eval and observability program run in CI as a merge gate: golden sets, LLM-as-judge scoring, span sampling, drift detection. The judge is a model too, so you evaluate the evaluator.
GxP validation for a system that never returns the same answer twice, and calibrated risk scoring where a conformal guarantee is marginal rather than conditional and the audit is not.
vLLM, SGLang and TensorRT-LLM at seven-figure inference scale, and an 8B assistant compressed to run fully offline on a mid-tier Android handset with no connectivity.
An extraction confidence is not an adverse-action reason. Document Intelligence returns two scores, and the gate reads the wrong one.
An HCM grants permissions to people holding jobs, and an agent is neither. Least privilege with no native expression, and the primitive that just shipped.
Four vendor agent platforms in one mesh, each sound alone. The seven joins that failed, starting with three catalogs where the loosest one decides.
The stuff that doesn't fit in a commit message — calls I'd revisit, what broke in production and why, and the parts I still find genuinely fun. Evaluation, retrieval, serving, the occasional war story.
The judge is a model too, with the same length, position and self-preference biases. Measuring it, versioning it, and knowing when not to use one.
Containment counts the sessions that never needed you. A booking agent's rollback is a refund, and a surge multiple is a scalar over a rotating mixture.
One roofline number sets the speculative-decoding gate, the batch-size knee and the chunked-prefill floor. Why max-num-seqs is a latency control.
Validating a system that never returns the same answer twice: what qualification becomes when the output is distributional, and where sampling breaks.
Reach me at mikitadaroshkin@gmail.com.