Specialties

Solutions

Contact us

Book demo

Contact us

Book demo

Specialties

Solutions

Contact us

Book demo

Human-agent execution model

Author: Mac Klinkachorn


Most "AI with a human in the loop" products mean one of two things. Either a person reviews everything the model produces, which caps throughput at the speed of the reviewer, or the model does everything and a person gets paged when a customer complains. Neither works for prior authorization or benefits verification. The first is too slow for tens of millions of prescriptions a year. The second puts a patient's therapy start date on the line.


The Trellis Agent Platform runs a third model. An agent picks up a task and tries to finish it. A harness checks the work. If the harness passes and the agent is confident, the task goes through without a human touching it. If not, the task moves to a human, who finishes it and hands it back. Every escalation becomes training signal. This post walks through that loop, stage by stage, and ties each stage to the research we lean on.



Stage 1: An agent picks up a task


Work enters through the Production layer. Production Inputs arrive from the pharmacy or infusion system (WellSky, CPR+, CareTend), from a payer portal, or from a fax queue. The Normalize step maps each one onto a workflow definition from the Workflow spec: a prior auth submission, an insurance discovery run, an appeal, a coverage determination. The Serving Handler assigns the normalized task to an agent runtime.


The agent runtime is the Tool loop. The model proposes an action, the runtime executes it, the result returns to the model, and the loop iterates. Each proposed action passes through the Safety layers before dispatch: HIPAA access checks, connector lifecycle, permissions, and audit logging. Approved actions dispatch to either the Tool registry (AI block, code interpreter, computer command) or the Connector registry (EHR group, claims system, clinical docs, payer API, taxonomy).


Tracing instruments every step. That matters later. The trace is not a log for debugging after the fact; it is the primary input to the harness.



Stage 2: The agent runs the task to completion


We deliberately let the agent attempt the whole task. It is tempting to insert human checkpoints at every risky step, but checkpoints fragment the trace, and a fragmented trace is hard to evaluate. Anthropic's guidance on [building effective agents](https://www.anthropic.com/research/building-effective-agents) draws the line between workflows (predefined code paths) and agents (the model directs its own process). Trellis runs both: the workflow definition fixes the control flow and state, and the agent fills in the actions inside each step. That hybrid keeps traces comparable across runs, which is what makes evaluation tractable.


The agent carries context from the Knowledge layer: historical denials, clinical SOPs, worked examples, prior auth cases, and payer guidelines. It does not get to invent policy. When a payer's criteria for a biologic require a documented step-therapy failure, that criterion lives in Knowledge, and the harness checks against it.


Long tasks are where agents break. [METR's time-horizon work](https://arxiv.org/abs/2503.14499) measures how long a task a model can complete at 50% reliability and found frontier models at roughly an hour in early 2025, doubling about every seven months. A full prior auth with portal navigation, document retrieval, and form entry sits near that edge. So the runtime caps attempts and routes failures to the Failure handler rather than letting the agent loop.


Stage 3: The harness evaluates the work


When the agent declares the task complete, the trace and its outputs go to the Evaluation substrate. The harness runs three tiers, and the order matters: cheap deterministic checks first, domain rules second, model judgment last.


Tier 1, deterministic harness. Did the agent produce every required artifact? Does the submitted form match the expected outputs for this workflow type? Are the NPI, member ID, NDC, and diagnosis codes well formed and consistent across documents? These checks are code, not models. They are fast and they never hallucinate.


Tier 2, healthcare-specific rules.This is the rule pack that an engineer without domain background cannot write. Does the clinical justification cite the payer's actual criteria for this drug and this plan year? Was the request submitted under the correct benefit (medical vs. pharmacy) for the site of care? Did the agent touch PHI outside the connector scope? Does the quantity match the prescribed regimen? Rules come from payer guidelines, clinical SOPs, and mentor feedback from our clinical team, and they version alongside the workflows.


Tier 3, trace review. A judge model reads the full trace and scores the reasoning diff and action diff against a known-good reference. We score the process, not just the final answer. [Let's Verify Step by Step](https://arxiv.org/abs/2305.20050) showed that step-level supervision beats outcome-only supervision on multi-step reasoning, and [Agent-as-a-Judge](https://arxiv.org/abs/2410.10934) extended that to agentic traces, matching human evaluators where plain LLM-as-a-judge did not. We also take the warnings seriously. [AgentRewardBench](https://arxiv.org/abs/2504.08942) found that no single judge model wins across task types, and [TRAIL](https://arxiv.org/abs/2505.08638) found the best long-context model localized errors in agent traces only 11% of the time. A judge model is a signal, not a verdict. That is why it sits behind two deterministic tiers rather than in front of them.


Stage 4: Pass-through requires both a passing harness and high confidence


Two signals have to agree before a task clears without a human. The harness has to pass on all three tiers, and the agent's calibrated confidence has to clear the threshold for that workflow type.


Confidence is not the model's self-reported number. We learned that lesson the same way the field did. OpenAI's [Why Language Models Hallucinate](https://arxiv.org/abs/2509.04664) makes the structural argument: models trained and scored on benchmarks that reward guessing learn to guess, so a raw "I'm 95% sure" is worth little. The abstention literature, summarized in [Know Your Limits](https://arxiv.org/abs/2407.18418), treats the ability to say "not sure, hand this off" as a capability to be measured and trained, not an afterthought.


Our confidence score is a Correctness Verifier output calibrated per workflow and per payer against historical outcomes: did tasks with this trace signature get approved, denied, or kicked back? Thresholds are tighter where the cost of being wrong is higher. A benefits check that returns a copay estimate can clear at a lower bar than an appeal that commits a clinical argument to a payer's record.




Stage 5: Escalation, human completion, and the handback


Anything that fails the gate goes to the Escalation Handler with a reason code: which harness tier failed, which rule fired, or which confidence threshold was missed. The reason code decides who picks it up. A formatting failure goes to the operations queue. A clinical criteria failure goes to a clinician. A connector failure goes to engineering.


The human works inside the same workflow state the agent left behind. They see the trace, the partial outputs, and the specific failure. They do not start over. They complete the task, and the handback records three things: what the human changed (the action diff), why they changed it (the reasoning diff, captured as structured mentor feedback), and whether the original harness verdict was right.


Then the task transfers back to an agent for the remaining steps: submitting the completed form, writing status back to the pharmacy system, scheduling the follow-up check. The human does the part that needed judgment and nothing else.


This is a dual-control problem, and it is harder than it looks. [τ²-bench](https://arxiv.org/abs/2506.07982) measured agents in environments where both the agent and a human modify shared state and found performance drops sharply compared with agent-only settings. Our handback protocol exists to shrink that drop: state is explicit, ownership is explicit, and the agent re-reads the workflow state on resume rather than assuming nothing moved.




Stage 6: Escalations feed the self-improving loop


An escalation is a labeled failure. The self-improving loop treats it that way. The Failure Analyzer takes the human's reasoning diff and action diff and classifies the root cause: missing knowledge, wrong action sequence, bad rule, or model error. The Patch generator proposes a fix at the right layer, which could be a new worked example in Knowledge, an updated justification template, a change to the workflow's control flow, or a new harness rule. The patch goes to the Versioned store, the Test Harness replays it against the clinical test cases, and the Output Eval decides whether to promote. Three failed attempts and the loop escalates to a clinician, whose answer feeds clinical knowledge directly.


The design borrows from the trajectory of self-improvement research without copying its risk profile. [Reflexion](https://arxiv.org/abs/2303.11366) showed that verbal self-feedback stored in memory improves later attempts. [Darwin Gödel Machine](https://arxiv.org/abs/2505.22954) went further and let agents rewrite their own code, moving SWE-bench performance from 20% to 50%, with every change validated by a benchmark before it is kept. The [self-evolving agents survey](https://arxiv.org/abs/2507.21046) frames the field around what evolves (model, memory, tools, workflow), when, and how. Trellis evolves the workflow, the knowledge, and the harness. It does not let an agent edit the safety layers or the healthcare rule pack. Those change through human review only.




Why the healthcare rule pack is the product


Generic agent benchmarks tell you how good a model is at being an agent. They do not tell you whether a prior auth will be approved. [OSWorld](https://arxiv.org/abs/2404.07972) reported humans at 72% and the best model at 12% on real computer tasks at launch; models have improved a great deal since, but GUI grounding is still where runs fail. [MedAgentBench](https://arxiv.org/abs/2501.14654) put agents in a FHIR environment with 300 clinically derived tasks and the best model cleared about 70%. [HealthBench](https://arxiv.org/abs/2505.08775) used 262 physicians and 48,562 rubric criteria to grade medical conversations. That is the right instinct: in healthcare, the rubric is the hard part.


That is our position too. The model will keep getting better and we will keep swapping it in. The harness, the healthcare rules, the calibration data, and the escalation reason codes are what make an agent safe to run against a real patient's coverage. No model release hands you those.


What this buys a pharmacy


Clean tasks clear in minutes with no human touch. Ambiguous tasks reach the right human with the trace already assembled and the failure already named. And every one of those human decisions makes the next clean task a little more likely. That is the loop.


---

References


- Anthropic. [Building effective agents](https://www.anthropic.com/research/building-effective-agents). December 2024.

- Lightman et al. [Let's Verify Step by Step](https://arxiv.org/abs/2305.20050). OpenAI, 2023.

- Shinn et al. [Reflexion: Language Agents with Verbal Reinforcement Learning](https://arxiv.org/abs/2303.11366). 2023.

- Zhuge et al. [Agent-as-a-Judge: Evaluate Agents with Agents](https://arxiv.org/abs/2410.10934). Meta AI, 2024.

- Xie et al. [OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments](https://arxiv.org/abs/2404.07972). 2024.

- Jiang et al. [MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents](https://arxiv.org/abs/2501.14654). Stanford, 2025.

- Kwa et al. [Measuring AI Ability to Complete Long Tasks](https://arxiv.org/abs/2503.14499). METR, 2025.

- Lù et al. [AgentRewardBench: Evaluating Automatic Evaluations of Web Agent Trajectories](https://arxiv.org/abs/2504.08942). 2025.

- Deshpande et al. [TRAIL: Trace Reasoning and Agentic Issue Localization](https://arxiv.org/abs/2505.08638). 2025.

- Arora et al. [HealthBench: Evaluating Large Language Models Towards Improved Human Health](https://arxiv.org/abs/2505.08775). OpenAI, 2025.

- Zhang et al. [Darwin Gödel Machine: Open-Ended Evolution of Self-Improving Agents](https://arxiv.org/abs/2505.22954). Sakana AI / UBC, 2025.

- Barres et al. [τ²-Bench: Evaluating Conversational Agents in a Dual-Control Environment](https://arxiv.org/abs/2506.07982). Sierra, 2025.

- Gao et al. [A Survey of Self-Evolving Agents](https://arxiv.org/abs/2507.21046). 2025.

- Wen et al. [Know Your Limits: A Survey of Abstention in Large Language Models](https://arxiv.org/abs/2407.18418). 2024.

- Kalai et al. [Why Language Models Hallucinate](https://arxiv.org/abs/2509.04664). OpenAI, 2025.

> Start today <

Let us take care of the MESS(y) data

Click below to get access to our product and explore how Trellis can help you and your organization unlock new insights!

Book demo

2026 Trellis. All rights reserved. Contact: support@runtrellis.com

2026 Trellis. All rights reserved. Contact: support@runtrellis.com