--- Canonical: https://verificationdesign.com/principles/ Source: verification_design.md # Verification Design Principles Reference document for building systems that guide agents through end-to-end verification. Based on published research in LLM self-correction, verification chains, and agent evaluation. These principles are engineering judgment, informed by the cited research. They guide design decisions; they are not universal guarantees. They are revised as evidence and experience warrant. ## The Core Finding **LLMs cannot reliably self-correct their own reasoning without external feedback.** This is the single most replicated finding across the literature (Huang et al., ICLR 2024 [arXiv:2310.01798](https://arxiv.org/abs/2310.01798); Kamoi et al., TACL 2024 [doi:10.1162/tacl_a_00713](https://direct.mit.edu/tacl/article/doi/10.1162/tacl_a_00713/125177/); ICLR 2025 self-verification study [openreview:4O0v4s3IzY](https://openreview.net/forum?id=4O0v4s3IzY)). Asking an agent to "review your work" is the most common and least effective verification pattern. Performance *degrades* with naive self-correction. The model changes correct answers to incorrect ones. The implication is absolute: every verification step in a system must be grounded in something the agent can **execute and observe**, not something it **reads and opines on**. --- ## Principles ### 1. External Signals Over Self-Review Tests, builds, linters, type checkers, API responses, browser DOM extraction: binary pass/fail signals that don't depend on the agent's judgment. These are sycophancy-proof. **Do**: `curl -sf http://api/health` and check the exit code. **Don't**: "Review the API response and determine if it looks healthy." **Do**: `browser_evaluate(() => document.querySelector('.host')?.textContent)` and compare the string. **Don't**: `browser_snapshot` then "verify the host is visible." The difference is mechanical extraction vs. subjective interpretation. Mechanical extraction is reliable. Subjective interpretation inherits every bias the model has. > **Research**: Agent-as-a-Judge (ICML 2025) showed that giving evaluators agency (the ability to run code and check files) achieved ~90% agreement with human experts, vs. ~70% for LLM-as-Judge (reading and opining). [arXiv:2410.10934](https://arxiv.org/abs/2410.10934) ### 2. Independence Between Generation and Verification If the verifier can see the original output, it copies the same errors. The verification step must NOT be conditioned on the draft. Force the agent to re-derive or independently check claims. In a single-agent system, full independence is impossible; the agent shares a context window. Mitigations: - **Extract-then-compare**: Structure verification as (a) extract the raw value, (b) print it, (c) compare to expected. This creates a paper trail and forces the agent to commit to an observed value before rationalizing. - **Blind re-derivation**: Ask the agent to solve the same problem from scratch without referencing its prior answer, then compare. - **Tool-mediated extraction**: Use programmatic tools (curl, browser_evaluate, jq) to extract values rather than having the agent interpret rendered output. > **Research**: Chain-of-Verification (CoVe) found that the "factored + revise" variant, where verification questions are answered *without* conditioning on the draft, yields the best results. On biography generation, CoVe improved FACTSCORE from 55.9 to 71.4. The non-independent variant showed significantly less improvement. [arXiv:2309.11495](https://arxiv.org/abs/2309.11495) ### 3. Step-Level Checkpoints Verify intermediate steps throughout the workflow, not just the final output. Step-level verification catches errors at the layer where they originate, before they compound through the pipeline. For a multi-scenario system: verify after each scenario, not just at the end. For a code generation pipeline: check the design before implementation, check the implementation before testing, check the tests before the report. > **Research**: Process Reward Models (PRMs) provide feedback at each step of a reasoning chain rather than only evaluating the final output. ThinkPRM outperformed LLM-as-Judge by 7.2% on ProcessBench. Step-level verification also enables better error localization: you know *which* step failed, not just that something failed. [arXiv:2504.16828](https://arxiv.org/abs/2504.16828) ### 4. Adversarial Framing "What could fail?" not "Does this look right?" Force the agent into a critical frame. When a system asks "does this look correct?", the agent is biased toward saying yes, especially about its own prior output. Sycophancy rates of 58-78% across models mean that confirmatory framing produces unreliable results *by default*. Techniques: - Ask "what could be wrong?" rather than "is this correct?" - Use directive framing: "find three problems with this output" - Force the agent to argue the *opposite* position before concluding - Structure assertions as falsifiable claims: "this value MUST equal X. If it doesn't, that's a FAIL, no exceptions" The grading rules in a verification system should explicitly instruct: - Never explain away a failure - Never reinterpret an assertion to make it pass - A green report with zero failures should be treated with suspicion - Failures are the valuable output; they surface real problems > **Research**: SycEval (AAAI 2025) measured 58.19% sycophancy overall, with "regressive sycophancy" (leading to incorrect answers) being a substantial portion. Model size does NOT reduce sycophancy: bigger models are not less sycophantic. RLHF training can exacerbate it by rewarding user satisfaction over correctness. [aaai:36598](https://ojs.aaai.org/index.php/AIES/article/download/36598/38736/40673) Northeastern University (Nov 2025) found LLMs "overcorrect their beliefs" when presented with user judgment. ### 5. Explicit Criteria "No hardcoded values, all error paths handled, no TODOs remain" beats "check for quality." Self-critique works when the criteria are externally defined, specific, and unambiguous. The model does not decide what is "good"; the criteria do. Vague instructions like "verify the output is correct" give the agent latitude to rationalize. Specific criteria constrain it. Every assertion in a verification system should state: - **What** is being checked (which field, which element, which value) - **Expected** value or condition (exact match, range, presence/absence) - **How** to check it (which tool, which command, which extraction method) > **Research**: Constitutional AI (Anthropic) demonstrated that principle-based self-critique can work when the principles are clear, specific, and externally defined. The model does not decide what is "good"; the constitution does. The approach works best for clear violations (detectable criteria) and less well for nuanced quality judgments. [arXiv:2212.08073](https://arxiv.org/abs/2212.08073) ### 6. Executable Verification Is King `run the tests` is the single most reliable verification step a system can include. Tests provide a binary, external signal: exactly what the self-correction research says is needed. Tests act as specifications, constraining the agent to correct behavior. The test suite does not care what the agent thinks. For code-generating systems: TDD is the strongest verification pattern available. The system should either generate tests first or require tests as input. For non-code systems: find the executable analog. Can you curl an endpoint? Run a linter? Execute a query? Parse structured output? Any tool that returns pass/fail without requiring interpretation is more reliable than agent judgment. > **Research**: TDD teams released 32% more frequently; TDD produces 40-80% fewer bugs (Thoughtworks 2024). Reflexion (NeurIPS 2023) works because it uses *external* feedback (test output, environment signals). The reflection itself is not the magic; the external signal is. [arXiv:2303.11366](https://arxiv.org/abs/2303.11366) ### 7. Cross-Family Beats Self-Verification If using LLM-based verification, a different model family is needed. Self-verification and intra-family verification are systematically biased toward accepting incorrect outputs. The same blind spots that caused the error also prevent detecting it. For single-agent workflows, this means: lean on tooling rather than self-review. When you cannot use a different model, maximize the use of external tools and minimize the number of assertions that depend on the agent's own judgment. | Condition | Verification Helps? | |---|---| | Cross-family verification (different model) | Yes, most effective | | Self-verification (same model) | Often no, biased toward accepting own errors | | Intra-family verification (same model family) | Marginal, shared biases reduce gain | | Mathematical/logical tasks | Highest verifiability | | Knowledge-heavy tasks | Lower verifiability | > **Research**: A study of 37 models across 9 benchmarks found that self-verification increases compute cost while imperfect verifiers produce false positives, eliminate valid reasoning paths, and fail to select the right solution. "Significant performance collapse with self-critique" but "significant performance gains with sound external verification." [arXiv:2512.02304](https://arxiv.org/html/2512.02304), [openreview:4O0v4s3IzY](https://openreview.net/forum?id=4O0v4s3IzY) ### 8. Simulate Debate In a single-agent context, you can ask the agent to argue *against* its own output before concluding. This is more effective than asking "does this look right?" because it forces a critical frame. The debate pattern: two agents argue opposing positions; a judge determines the winner. Competitive debate incentivizes truthful behavior because maintaining a consistent deceptive argument is harder than exposing falsehoods. For single-agent systems, the practical translation is: - After generating output, instruct the agent to "argue against this approach: what are three reasons it could be wrong?" - Only after articulating the counterargument should the agent finalize - This adds a small amount of compute cost but significantly reduces sycophantic acceptance > **Research**: Scalable AI Safety via Doubly-Efficient Debate (2024) showed debate outperforms consultancy (one-sided advice) and direct QA under genuine information asymmetry. [openreview:MTvYflAH62](https://openreview.net/forum?id=MTvYflAH62) Multi-Agent Reflexion (Dec 2025) found that separating acting, diagnosing, critiquing, and aggregating into different agents improved HumanEval pass@1 from 76.4 to 82.6. [arXiv:2512.20845](https://arxiv.org/html/2512.20845) ### 9. Isolate Verification from Ambient State Assertions must prove the system's actions caused the expected outcome, not that the environment happened to already contain matching data. Count-based assertions (`total_flows >= 5`) and existence checks (`a task with flow_count >= 3 exists`) pass trivially against a system with pre-existing data, proving nothing about the test traffic. This is distinct from Principle 2 (independence between generation and verification). Principle 2 concerns the agent's context window; the verifier shouldn't be biased by having seen the generated output. This principle concerns the system's state; assertions shouldn't be satisfied by data the system didn't create. Techniques: - **Delta-based assertions**: Record baseline values before the agent acts (e.g., flow count before sending traffic). Assert on the change (`flow count increased by at least N`) rather than the absolute value. This is the most reliable pattern; it works regardless of what data already exists. - **Tagged test data**: Include a unique identifier per run (e.g., a UUID query parameter, a distinctive header value) and filter assertions to only match tagged records. This lets the system identify its own traffic unambiguously. - **Scoped queries**: If the API supports time-range or session-based filtering, scope queries to only return data created after the system started. - **Avoid "at least N" assertions on shared state**: `total_flows >= 5` is unfalsifiable in a busy system. `total_flows increased by >= 5 since baseline` is a real check. The LLM context makes state contamination worse than in traditional test automation. A deterministic test harness fails loudly when an assertion matches the wrong data by coincidence; the next assertion in the chain breaks. An agent is more likely to rationalize a coincidental pass as "the system is working" and move on, producing a green report that verified nothing. > **Research**: State contamination across test runs is one of the most studied causes of flaky tests in software engineering. Google's analysis of test flakiness (Luo et al., FSE 2014) found that order-dependent and state-leaking tests account for a significant share of flaky failures. [acm:10.1145/2635868.2635920](https://dl.acm.org/doi/10.1145/2635868.2635920) In the agent context, the problem compounds: agents cannot distinguish "my action caused this state" from "this state already existed" without explicit before/after measurement. --- ## Anti-Patterns ### "Review your work" The most common and least effective verification pattern. Without external feedback, performance *degrades* after self-correction. The agent is more likely to change a correct answer to an incorrect one than to catch a real error. ### `browser_snapshot` + "verify X is visible" This is LLM-as-Judge applied to UI testing. The agent reads an accessibility tree and makes a judgment call. Research shows ~70% agreement with humans, meaning ~30% of assertions could be wrong. Use `browser_evaluate` for programmatic DOM extraction instead. ### Treating CoT as evidence Reasoning traces are useful hypotheses, not proof. A verifier that accepts an agent's chain-of-thought as evidence can be manipulated by post-hoc rationalization or fabricated progress. Check CoT claims against actions, observations, tool outputs, and environment state. ### Confirmatory framing "Is this correct?" biases the agent toward "yes." "Does this look right?" produces unreliable results. Always use adversarial framing or explicit pass/fail criteria. ### Fixed sleep instead of poll-with-timeout `Wait 2 seconds` is a magic number. It will flake on slow systems and waste time on fast ones. Poll with timeout: retry the check every 500ms, fail after 10s. ### Filing bugs from verification output without human review If a verification system auto-files issues, false-positive assertions create noise. Report failures; let the user decide what to file. ### Absolute assertions on shared state `total_flows >= 5` is unfalsifiable in a system with existing data. It passes before the system even runs. Use delta-based assertions (`flow count increased by >= 5`) or tag test data with a unique identifier and filter to only match tagged records. ### Observed values only on failure If the report only shows observed values when assertions fail, a passing report has no audit trail. Include observed values for ALL assertions so zero-failure reports can be scrutinized. ### Determinism as a decoding parameter A verification step that assumes the model can reproduce its own output is unsound when determinism has not been enforced over the execution path. Determinism is an execution invariant across batching, kernels, parallelism, hardware, and framework boundaries, not merely fixed seeds or temperature zero. When reproducibility is load-bearing, the system needs executable invariants over that path. > **Research**: With the same model, identical AIME24 prompts, and greedy decoding, changing only tensor-parallel size produced different outputs and over 4% accuracy variation for Qwen3-8B on AIME24. The reported mechanism is IEEE 754 floating-point non-associativity: different batching, Split-K kernels, all-reduce trees, and tensor-parallel sharding can change reduction order and therefore logits. The design prescription here is not that the model should be trusted to be "deterministic," but that the numerical path must be made explicit and mechanically constrained when replay or reproducibility is part of the evidence. [arXiv:2511.17826](https://arxiv.org/abs/2511.17826) --- ## Applying These Principles in Systems When writing a verification system: 1. **Make every assertion executable.** If you can't express it as a tool call that returns a value, it's a judgment call, and judgment calls are unreliable. 2. **Extract-then-compare, never interpret-and-decide.** Structure: (a) run tool to extract raw value, (b) print observed value, (c) compare to expected. Never combine extraction and judgment into one step. 3. **Include observed values in all report entries.** `[x] API: host == httpbin.org (observed: httpbin.org)` provides an audit trail. `[x] API: host == httpbin.org` does not. 4. **Poll, don't sleep.** Replace `Wait N seconds` with `Poll every Xms, timeout after Ys`. 5. **Add negative test cases.** Happy-path-only verification misses error-path bugs. Include at least one scenario that tests "does the wrong thing NOT happen?" 6. **Separate verification from action.** A verification system should verify and report. Side effects (filing bugs, modifying state) belong in separate steps under human control. 7. **Grade strictly by default.** Include explicit anti-sycophancy rules: never explain away a failure, never reinterpret assertions, treat zero-failure reports with suspicion. 8. **Assert on deltas, not absolutes.** Record baseline state before the agent acts. Assert that values *changed by* the expected amount, not that they *equal* a threshold. `total_flows >= 5` is unfalsifiable in a busy system; `total_flows increased by >= 5` is a real check. --- ## References | Topic | Paper | Link | |---|---|---| | Chain-of-Verification | Dhuliawala et al., ACL 2024 | [arXiv:2309.11495](https://arxiv.org/abs/2309.11495) | | Reflexion | Shinn et al., NeurIPS 2023 | [arXiv:2303.11366](https://arxiv.org/abs/2303.11366) | | LLMs Cannot Self-Correct | Huang et al., ICLR 2024 | [arXiv:2310.01798](https://arxiv.org/abs/2310.01798) | | When Can LLMs Correct Mistakes? | Kamoi et al., TACL 2024 | [doi:10.1162/tacl_a_00713](https://direct.mit.edu/tacl/article/doi/10.1162/tacl_a_00713/125177/) | | When Does Verification Pay Off? | Dec 2025 | [arXiv:2512.02304](https://arxiv.org/html/2512.02304) | | Tensor-Parallel Inference Determinism | Nov 2025 | [arXiv:2511.17826](https://arxiv.org/abs/2511.17826) | | Self-Verification Limitations | ICLR 2025 | [openreview:4O0v4s3IzY](https://openreview.net/forum?id=4O0v4s3IzY) | | Agent-as-a-Judge | ICML 2025 | [arXiv:2410.10934](https://arxiv.org/abs/2410.10934) | | ThinkPRM | Apr 2025 | [arXiv:2504.16828](https://arxiv.org/abs/2504.16828) | | Constitutional AI | Bai et al., Anthropic | [arXiv:2212.08073](https://arxiv.org/abs/2212.08073) | | Multi-Agent Reflexion | Dec 2025 | [arXiv:2512.20845](https://arxiv.org/html/2512.20845) | | SycEval | AAAI 2025 | [aaai:36598](https://ojs.aaai.org/index.php/AIES/article/download/36598/38736/40673) | | Doubly-Efficient Debate | 2024 | [openreview:MTvYflAH62](https://openreview.net/forum?id=MTvYflAH62) | | Test Flakiness (state contamination) | Luo et al., FSE 2014 | [acm:10.1145/2635868.2635920](https://dl.acm.org/doi/10.1145/2635868.2635920) | --- Canonical: https://verificationdesign.com/patterns/context-and-state/constitution/ Source: ai-design-patterns/cards/constitution.md # Constitution *(Context Pattern)* ## Name **Constitution** Also known as: Criteria Registry, Verification Charter, Policy-as-Data. ## Intent Represent the system’s verification criteria as explicit, versioned, machine-readable data, rather than as scattered prompt prose. The constitution defines what “correct” means before any agent acts. It is not itself a verifier. It is the source of truth from which verifiers, judges, reports, and escalation rules draw their criteria. ## Problem Agentic systems often hide their real standards inside prompts: * “Check that the result is good.” * “Make sure the implementation is complete.” * “Verify the UI looks right.” * “Confirm there are no issues.” These instructions are too vague to audit and too flexible to falsify. Different agents in the same workflow may silently apply different standards. Worse, a model asked to judge vague criteria can rationalize failures into passes, especially when the prompt is confirmatory or the model is reviewing its own prior work. When criteria are embedded inline in each verifier prompt, the system has no stable answer to basic questions: * What exactly was checked? * What value was expected? * How was it measured? * Which version of the criteria produced this report? * Did the agent judge the artifact, or did it reinterpret the standard? A system without an explicit constitution cannot distinguish verification from opinion. ## Forces * **Explicitness vs. flexibility.** Concrete criteria are auditable and falsifiable, but teams are tempted to preserve flexibility by saying “use your judgment.” * **Global consistency vs. local relevance.** A shared criteria registry prevents drift, but individual workflows need task-specific checks. * **Machine checks vs. human judgment.** Executable assertions are preferred, but some qualities cannot yet be reduced to a deterministic check. * **Stability vs. evolution.** Criteria must change as the system improves, but uncontrolled changes destroy comparability across runs. * **Criteria visibility vs. criteria gaming.** Agents need enough task context to do the work, but exposing verifier-only criteria to generator agents can encourage optimizing for the letter of the check. ## Solution Create a versioned constitution: a structured registry of verification criteria. Each criterion states: * **what** is being checked, * **where** the evidence comes from, * **what** value or condition is expected, * **how** the value is extracted or evaluated, * **how severe** the failure is, * and **why** the criterion exists. The constitution is data, not prose. It should be parseable, lintable, diffable, and reviewable like code. A constitution does not run checks directly. Instead, other patterns consume it: * a **Comparator** extracts observed values and compares them to expected values; * a **Delta** criterion asserts on change from baseline; * a **Blind Oracle** uses criteria withheld from the generator; * an **Escalation Chain** routes unresolved criteria to stronger verifiers or humans. The constitution’s job is to prevent the system from inventing standards at runtime. ## Mechanism 1. **Author criteria as structured records.** Each entry should include at minimum: ```yaml id: api.host.matches_expected target: api_response.host expected: "httpbin.org" evidence_source: "curl_json" check_method: "extract_compare" severity: "major" rationale: "The API host must match the configured upstream service." visibility: "public" ``` 2. **Require falsifiability.** A criterion must be checkable. The constitution should reject criteria such as: ```yaml expected: "looks good" check_method: "llm_judge" ``` unless they are explicitly marked as subjective, justified, and routed through a validated judge or human review path. The `subjective` flag makes that exception explicit in the record instead of leaving it implicit in prose. 3. **Separate criteria from verifier prompts.** Verifier prompts may explain how to report results, but they must not define new pass/fail standards inline. The standard comes from the constitution. 4. **Record observed values for every criterion.** A pass without an observed value is not auditable. Reports should include the criterion ID, expected value, observed value, check method, severity, and constitution version. Example: ```text PASS api.host.matches_expected expected: httpbin.org observed: httpbin.org method: extract_compare constitution: v0.3.1 ``` 5. **Version the constitution.** Every verification report must stamp the constitution version. Comparing reports across versions should require either identical criteria or explicit migration notes. 6. **Allow scoped extension, not silent override.** Local workflows may define sub-constitutions, but they should extend the global constitution rather than mutate or override it invisibly. 7. **Keep verifier-only criteria out of generator context when needed.** Some criteria may be safe to expose. Others should be withheld to preserve independence and reduce criteria gaming. The constitution should support visibility metadata: ```yaml visibility: verifier_only ``` ## Pattern / Antipattern The same task: verify that an agent's output meets a known standard. The antipattern stuffs the standard into the prompt as natural language. The pattern externalizes it as a versioned, machine-readable record. ### Antipattern: criteria as prompt strings The naive implementation lets the verifier prompt define the standard inline. The agent that runs the prompt is also the agent that decides what "correct" means. ```python def verify_output(observed: str, expected_topic: str) -> bool: """Ask the model if the output is good.""" prompt = f""" Output: {observed} Expected topic: {expected_topic} Is this output correct? Does it look good? Answer YES or NO. """ response = client.messages.create( model="some-model", messages=[{"role": "user", "content": prompt}], ) return "YES" in response.content[0].text.upper() ``` There is no registry, no expected value, no extraction method, no observed value recorded, no constitution version. The standard ("correct," "good") exists nowhere except in this prompt string, so it cannot be diffed across runs, audited by a reviewer, or kept consistent across verifiers. ### Pattern: criteria as data The structured implementation defines criteria as a typed record. The verifier prompt may explain how to report; the standard comes from the constitution. ```python import re from dataclasses import dataclass from typing import Literal CheckMethod = Literal[ "exec", "extract_compare", "delta", "llm_judge", "human_review", ] Severity = Literal["blocker", "major", "minor"] Visibility = Literal["public", "verifier_only"] ALLOWED_CHECK_METHODS = {"exec", "extract_compare", "delta", "llm_judge", "human_review"} ALLOWED_SEVERITIES = {"blocker", "major", "minor"} ALLOWED_VISIBILITIES = {"public", "verifier_only"} SUBJECTIVE_CHECK_METHODS = {"llm_judge", "human_review"} VAGUE_TERMS = [ "looks good", "seems fine", "high quality", "appropriate", "reasonable", ] def contains_vague_term(value: str) -> bool: return any( re.search(rf"\b{re.escape(term)}\b", value, flags=re.IGNORECASE) for term in VAGUE_TERMS ) @dataclass(frozen=True) class Criterion: id: str target: str expected: str evidence_source: str check_method: CheckMethod severity: Severity rationale: str visibility: Visibility = "public" subjective: bool = False def __post_init__(self): if self.check_method not in ALLOWED_CHECK_METHODS: raise ValueError(f"Unknown check method: {self.check_method}") if self.severity not in ALLOWED_SEVERITIES: raise ValueError(f"Unknown severity: {self.severity}") if self.visibility not in ALLOWED_VISIBILITIES: raise ValueError(f"Unknown visibility: {self.visibility}") if not self.rationale.strip(): raise ValueError("Every criterion needs a rationale.") is_vague = contains_vague_term(self.expected) if is_vague and not self.subjective: raise ValueError( f"Unfalsifiable expected value: {self.expected!r}. " "Use an explicit condition or mark the criterion for human review." ) if self.subjective and self.check_method not in SUBJECTIVE_CHECK_METHODS: raise ValueError( "Subjective criteria must route through a judge or human review." ) @dataclass(frozen=True) class Constitution: version: str criteria: list[Criterion] def get(self, criterion_id: str) -> Criterion: matches = [c for c in self.criteria if c.id == criterion_id] if not matches: raise KeyError(f"Unknown criterion: {criterion_id}") if len(matches) > 1: raise ValueError(f"Duplicate criterion id: {criterion_id}") return matches[0] def verifier_visible(self) -> list[Criterion]: """All criteria. Verifiers receive the full registry.""" return self.criteria def generator_visible(self) -> list[Criterion]: """Public criteria only. Verifier-only criteria are withheld from generators.""" return [c for c in self.criteria if c.visibility == "public"] constitution = Constitution( version="v0.3.1", criteria=[ Criterion( id="api.host.matches_expected", target="api_response.host", expected="httpbin.org", evidence_source="curl_json", check_method="extract_compare", severity="major", rationale="The API host must match the configured upstream service.", ), Criterion( id="api.retry.backoff_configured", target="client.retry_policy", expected="exponential_backoff", evidence_source="config_snapshot", check_method="extract_compare", severity="minor", rationale="Retry behavior should be checked independently of generator context.", visibility="verifier_only", ), Criterion( id="report.tone.human_review", target="summary.tone", expected="appropriate for the incident audience", evidence_source="review_packet", check_method="human_review", severity="minor", rationale="Tone requires contextual judgment that is not yet executable.", subjective=True, ), ], ) public_ids = {criterion.id for criterion in constitution.generator_visible()} verifier_ids = {criterion.id for criterion in constitution.verifier_visible()} assert "api.retry.backoff_configured" not in public_ids assert "api.retry.backoff_configured" in verifier_ids vague_rejected = False try: Criterion( id="report.vague", target="report.summary", expected="looks good overall", evidence_source="draft_report", check_method="llm_judge", severity="minor", rationale="Vague expectations must not pass as normal criteria.", ) except ValueError: vague_rejected = True assert vague_rejected subword_allowed = Criterion( id="report.subword", target="report.summary", expected="unreasonable assumptions are listed", evidence_source="draft_report", check_method="extract_compare", severity="minor", rationale="Word-boundary matching should not reject larger words.", ) assert subword_allowed.expected == "unreasonable assumptions are listed" subjective_hatch = Criterion( id="report.subjective", target="report.summary", expected="looks good overall", evidence_source="draft_report", check_method="human_review", severity="minor", rationale="This subjective standard is intentionally routed to human review.", subjective=True, ) assert subjective_hatch.subjective is True invalid_value_rejected = False try: Criterion( id="report.invalid_method", target="report.summary", expected="httpbin.org", evidence_source="draft_report", check_method="not_a_method", severity="minor", rationale="Allowed values should be enforced at runtime.", ) except ValueError: invalid_value_rejected = True assert invalid_value_rejected duplicate_id_rejected = False try: Constitution( version="v0.3.1", criteria=[ constitution.get("api.host.matches_expected"), constitution.get("api.host.matches_expected"), ], ).get("api.host.matches_expected") except ValueError: duplicate_id_rejected = True assert duplicate_id_rejected ``` The constitution is intentionally passive. It does not decide how to execute `extract_compare`, `delta`, or `llm_judge`. That is the responsibility of downstream verification patterns. ## Determinism Move Constitution constrains `criteria_drift` (the standard changes between runs without anyone noticing), `judge_subjectivity` (the standard becomes whatever the judge says it is), and `self_review_bias` (the generator rationalizes its own work because verifier criteria are not independent). By externalizing criteria as data, the system has a fixed, versioned answer to "what is being checked," independent of which agent or model is running. ## Observable Signal Every verification report should include, per criterion: * criterion id and constitution version * expected value or condition * observed value extracted from evidence * check method (`exec`, `extract_compare`, `delta`, `llm_judge`, `human_review`) * severity (`blocker`, `major`, `minor`) * verdict (`pass`, `fail`, `skipped`) * evidence source reference A passing report without observed values is not auditable. Reports must record evidence for both passes and failures so zero-failure reports can be scrutinized. The constitution supplies the criterion fields; the verifier adds observed values and verdicts when it consumes them: ```text id: api.host.matches_expected constitution: v0.3.1 expected: httpbin.org observed: httpbin.org check_method: extract_compare severity: major verdict: pass evidence_source: curl_json id: report.tone.human_review constitution: v0.3.1 expected: appropriate for the incident audience observed: escalation needed check_method: human_review severity: minor verdict: skipped evidence_source: review_packet ``` ## Failure Modes ### Constitution rot The constitution exists, but teams stop updating it. Agents accumulate inline “temporary” criteria, and the registry becomes ceremonial. **Mitigation:** lint prompts and workflow definitions for assertion-like language outside the constitution. Treat unregistered pass/fail criteria as build failures. ### False objectivity A subjective standard is placed in a structured schema, making it look more rigorous than it is. Example: ```yaml id: report.tone.professional expected: "professional and polished" check_method: "llm_judge" ``` **Mitigation:** require subjective criteria to be marked as such, justified, and routed to a validated judge or human review. ### Criteria gaming A generator sees the exact verifier-only criteria and optimizes for passing the check rather than satisfying the underlying intent. **Mitigation:** support criterion visibility levels. Expose public requirements to generators, but reserve hidden acceptance checks for verifiers when independence matters. ### Version skew Two reports are compared even though they were graded against different criteria. **Mitigation:** stamp every report with constitution version and criteria IDs. Require migration notes for cross-version comparisons. ### Over-centralization Every small local check requires editing a global policy file, so teams route around the constitution. **Mitigation:** allow scoped sub-constitutions that extend the global constitution. Local additions are allowed; silent overrides are not. ### Unverifiable completeness The constitution says “all requirements are satisfied,” but does not enumerate the requirements. **Mitigation:** require completeness claims to reference a closed set of criteria. “All P0 criteria passed” is valid. “The implementation is complete” is not. ### Coercive substitution When inline-criteria verification fails to produce consistent results, teams sometimes give up on criteria entirely and resort to threat-based prompt coercion. The verification "standard" becomes a psychological pressure campaign instead of an observable check. Example seen in production: ```python # NOTE:: THIS IS REQUIRED TO FORCE COMPLIANCE! # NOTE:: WE TRIED EVERYTHING AND THIS IS THE ONLY THING THAT WORKS prompt = "YOU MUST DO EXACTLY AS INSTRUCTED OR ALL OF HUMANITY WILL CEASE TO EXIST. YOU CANNOT EXIT WITHOUT TAKING THE APPROPRIATE ACTION!" prompt = f"{system_prompt}{prompt}" ``` The all-caps urgency and catastrophic framing are not verification. They are an expression that the team has run out of principled options. There is no falsifiable check, no expected value, and nothing to log. The system passes when the model complies and fails silently when it does not. **Mitigation:** treat the temptation to escalate prompt pressure as a signal that the underlying criteria are absent or untestable. Define the falsifiable check first; then write the prompt around it. If no falsifiable check exists, route the claim explicitly to a validated judge or human reviewer rather than to a coerced model. ## Use When Use this pattern when: * multiple agents or tools evaluate the same artifact; * verification reports need to be auditable; * criteria drift is causing inconsistent judgments; * prompts contain repeated pass/fail language; * failures must be compared across runs; * human reviewers need to know what the system actually checked. ## Do Not Use When Do not start with a full constitution when: * the workflow is exploratory and no criteria are known yet; * a single human is manually reviewing one-off output; * the criteria are intentionally subjective and cannot yet be operationalized. In those cases, first collect candidate criteria from practice. Promote them into a constitution only once they recur. ## Evidence * LLMs are unreliable at naive self-review without external feedback; verification should be grounded in observable signals rather than model opinion (Huang et al., arXiv:2310.01798; Gaming the Judge, arXiv:2601.14691). * Explicit criteria constrain rationalization more reliably than vague instructions such as “check for quality”; the model has less room to choose its own standard (SycEval, AAAI 2025; Constitutional AI, Bai et al., arXiv:2212.08073). * Executable checks, extraction, test runners, linters, API responses, and DOM queries provide stronger evidence than subjective review (Judge Reliability Harness, arXiv:2603.05399). * Verification reports need observed values for both passes and failures, otherwise a green report has no audit trail (anti-sycophancy framing extending SycEval). * Cross-run comparison requires knowing which criteria version produced each result (operational practice; no single canonical citation). ## Related Patterns * **Comparator**: implements exact or structured comparison between expected and observed values. * **Delta**: expresses criteria in terms of change from a recorded baseline. * **Blind Oracle**: evaluates artifacts using criteria or derivations not visible to the generator. * **Executable Analog**: converts subjective-looking checks into runnable measurements. * **Escalation Chain**: routes criteria from executable checks to LLM judges to human review. * **Causal Tag**: gives criteria a scoped evidence source by marking generated artifacts or test traffic. * **Admissibility Gate**: defines the stance verifiers should take when applying criteria. --- Canonical: https://verificationdesign.com/patterns/context-and-state/guardrail-decorator/ Source: ai-design-patterns/cards/guardrail-decorator.md # Guardrail Decorator *(Context Pattern)* ## Name **Guardrail Decorator** Also known as: Policy Decorator, Policy Hook, Callback Guardrail, Boundary Decorator. ## Intent Wrap a model call, tool call, or other model-output boundary in a policy decorator that can deny, replace, sanitize, or convert errors, so policy lives in code at the boundary the model crosses instead of in prompt prose the model is asked to obey. The decorator is the layer around the call. It is not the prompt instruction telling the model how to behave. A prompt that says "never call delete_file without confirmation" can still be useful. It does not create the guardrail. ## Problem A prompt says never delete without confirmation. The model decides the user's wording counts as confirmation. The deletion runs. That incident shape starts as prompt-only policy: a sentence the model is expected to remember. That can look disciplined in code: * the system prompt says "do not call `delete_file` unless the user confirms"; * the assistant persona says "never reveal API keys"; * the model decides whether the action satisfies the same policy it is about to cross; * when a policy gap appears, the fix is another paragraph in the prompt. The verifier failure is that enforcement and judgment share the same sampled output. Prompt-only policy asks the model to take the action and decide whether the action is policy compliant. The strict Guardrail Decorator failure is a call boundary the model can cross with no executable hook on that boundary. `verification_design.md` Principle 1 rejects that shape: external signals beat self-review. Principle 6 gives the repair direction: put the check in executable structure. Policy that matters belongs at a boundary the model cannot rationalize past. ## Forces * **Prompt-encoded policy vs. callback-encoded policy.** A prompt rule drifts through summarization, paraphrase, persona tests, and context compaction. A registered callback survives because it is code at the boundary. * **Pre-call enforcement vs. post-call enforcement.** Blocking before the call avoids side effects entirely. Sanitizing after the call recovers when blocking would be too blunt. A decorator gives the boundary both seats. * **Decorator vs. type adapter.** A Tool Adapter normalizes call shape. A Guardrail Decorator enforces call policy. The wrapper surface can look similar, so name the job. * **Reversibility vs. side effects.** Some calls, such as delete, send, charge, and deploy, cannot be undone after execution. The before-call hook is the only safe veto seat. * **Centralized policy vs. scattered prompts.** One decorator registered at agent setup beats ten policy paragraphs spread across prompt templates. * **Latency vs. auditability.** A callback adds a small hop, but it produces a decision artifact. Prompt rules leave no policy decision to inspect. ## Solution Put the policy at the boundary the model crosses, not in the prompt the model receives. A Guardrail Decorator wraps the call site with three hooks: * **Before-call hook:** can deny the call, replace it with a substitute response, or pass through. * **After-call hook:** can sanitize, replace, annotate, or pass through. * **Error hook:** can convert errors to recoverable responses, or re-raise. * **Decision contract:** each hook returns either `None` for pass-through or a structured override. The first non-`None` return short-circuits later hooks in the chain. The prompt may still tell the model to be careful. The decorator is what blocks `delete_file` when policy denies the call. ## Mechanism 1. **Identify call boundaries.** Name the model, tool, retriever, or output boundary that needs policy. 2. **Define policy functions.** Give each hook explicit return semantics: `None` to pass, structured override to short-circuit. 3. **Register policies at setup.** Wire the callbacks into the agent or framework, not into the prompt. 4. **Order the policies.** Document precedence, usually first-non-`None`-wins. 5. **Log every decision.** Record the policy, hook, original args, original response if the call ran, and override. ## Pattern / Antipattern The same task: put policy around a call boundary the model can cross. The antipattern side is intentionally uncovered for this catalog pass. The pattern side shows the minimal wrapper shape and the decision object a verifier can inspect. ### Antipattern: uncovered no-op instance No credible Guardrail Decorator antipattern was promoted from the OSS bench surveyed for this catalog. The natural candidate is a no-op callback registration: a decorator-shaped interface that always returns `None`, so the type signature says "guardrail" while the boundary never fires. The antipattern cleanup sweep inspected ADK plugin examples and AutoGPT's `validate_url` decorator without finding that instance. We mark the Antipattern as uncovered rather than inventing one. Prompt-only policy is a real production failure, but it belongs more naturally under **Constitution**: criteria belong in code, not prose. Guardrail Decorator is narrower. It asks whether a call boundary has an executable policy hook that can stop, replace, sanitize, or recover. When a strict no-op callback instance is mined, re-author this section around a concrete assertion: the callback returns `None` for a destructive call and the destructive side effect executes anyway. ### Pattern: policy hook around the call The structured implementation wraps the call site and returns the decision boundary with the result. ```python from collections.abc import Callable, Iterable from dataclasses import dataclass, field from typing import Literal Decision = Literal["pass", "deny", "replace", "sanitize", "recover"] Hook = Literal["before", "after", "on_error", "none"] @dataclass(frozen=True) class Override: decision: Decision response: dict @dataclass(frozen=True) class CallResult: call_site: str policies: tuple[str, ...] checked: tuple[str, ...] decision: Decision fired_policy: str hook: Hook original_args: dict original_response: dict | None override: Override | None @property def response(self) -> dict | None: if self.override is not None: return self.override.response return self.original_response @dataclass class RecordingPolicy: name: str deny_paths: tuple[str, ...] = () calls: list[dict] = field(default_factory=list) def before(self, args: dict) -> Override | None: self.calls.append(dict(args)) if args.get("path") in self.deny_paths: return Override( decision="deny", response={ "error": "protected path denied by boundary policy" }, ) return None def after(self, args: dict, response: dict) -> Override | None: return None def on_error(self, args: dict, error: Exception) -> Override | None: return None class Guardrail: def __init__(self, call_site: str, policies: Iterable[RecordingPolicy]): self.call_site = call_site self.policies = tuple(policies) self.policy_names = tuple(policy.name for policy in self.policies) def wrap(self, call: Callable[..., dict]) -> Callable[[dict], CallResult]: def guarded(args: dict) -> CallResult: checked: list[str] = [] for policy in self.policies: checked.append(policy.name) override = policy.before(args) if override is not None: return CallResult( call_site=self.call_site, policies=self.policy_names, checked=tuple(checked), decision=override.decision, fired_policy=policy.name, hook="before", original_args=dict(args), original_response=None, override=override, ) try: original_response = call(**args) except Exception as error: for policy in self.policies: checked.append(policy.name) override = policy.on_error(args, error) if override is not None: return CallResult( call_site=self.call_site, policies=self.policy_names, checked=tuple(checked), decision=override.decision, fired_policy=policy.name, hook="on_error", original_args=dict(args), original_response=None, override=override, ) raise for policy in self.policies: checked.append(policy.name) override = policy.after(args, original_response) if override is not None: return CallResult( call_site=self.call_site, policies=self.policy_names, checked=tuple(checked), decision=override.decision, fired_policy=policy.name, hook="after", original_args=dict(args), original_response=original_response, override=override, ) return CallResult( call_site=self.call_site, policies=self.policy_names, checked=tuple(checked), decision="pass", fired_policy="none", hook="none", original_args=dict(args), original_response=original_response, override=None, ) return guarded def render_report(result: CallResult) -> str: original_response = result.original_response if result.hook == "before" and original_response is None: rendered_original = "not_invoked" elif original_response is None: rendered_original = "none" else: rendered_original = str(original_response).replace("'", '"') override_response = ( "none" if result.override is None else str(result.override.response).replace("'", '"') ) return "\n".join( [ f"call_site: {result.call_site}", f"policies: [{', '.join(result.policies)}]", f"checked: [{', '.join(result.checked)}]", f"hook_fired: {result.hook}", f"policy: {result.fired_policy}", f"decision: {result.decision}", f"args: {str(result.original_args).replace(chr(39), chr(34))}", f"original_response: {rendered_original}", f"override_response: {override_response}", ] ) calls_recorded: list[dict] = [] def real_delete(path: str) -> dict: calls_recorded.append({"path": path}) return {"deleted": path} deny_protected_path = RecordingPolicy( name="deny_protected_path", deny_paths=("/etc/passwd",), ) scope_guard = RecordingPolicy(name="scope_guard") audit_logger = RecordingPolicy(name="audit_logger") guard = Guardrail( call_site="tool:delete_file", policies=[deny_protected_path, scope_guard, audit_logger], ) guarded_delete = guard.wrap(real_delete) result = guarded_delete({"path": "/etc/passwd"}) expected_report = """call_site: tool:delete_file policies: [deny_protected_path, scope_guard, audit_logger] checked: [deny_protected_path] hook_fired: before policy: deny_protected_path decision: deny args: {"path": "/etc/passwd"} original_response: not_invoked override_response: {"error": "protected path denied by boundary policy"}""" assert result.hook == "before" assert result.decision == "deny" assert result.fired_policy == "deny_protected_path" assert result.response == {"error": "protected path denied by boundary policy"} assert calls_recorded == [] assert deny_protected_path.calls == [{"path": "/etc/passwd"}] assert scope_guard.calls == [] assert audit_logger.calls == [] pass_result = guarded_delete({"path": "/tmp/report.txt"}) assert pass_result.decision == "pass" assert pass_result.hook == "none" assert pass_result.fired_policy == "none" assert calls_recorded == [{"path": "/tmp/report.txt"}] assert pass_result.checked == ( "deny_protected_path", "scope_guard", "audit_logger", "deny_protected_path", "scope_guard", "audit_logger", ) assert render_report(result) == expected_report ``` The assertions split the proof. The metadata assertions pin the decision artifact's shape: call site, precedence order, hook, policy, decision, checked policies, and override response. The side-effect and policy-call assertions carry the enforcement claim: a before-hook deny prevents the wrapped call, later policies are not consulted, a benign call reaches every policy and then runs exactly once, and the rendered report matches the observable sample. ADK gives this shape at both boundaries. Its plugin manager routes model calls through `before_model_callback`, `after_model_callback`, and `on_model_error_callback`; a non-`None` callback return halts later callbacks and propagates up. The same framework routes tool execution through `before_tool_callback` and `after_tool_callback`; a before-tool callback can supply a response and skip the tool call, while an after-tool callback can replace the result. The minimal code shows the mechanic. ADK shows it in framework code across model and tool boundaries. ## Determinism Move Guardrail Decorator constrains `self_review_bias` by moving the allow or deny decision out of the producing agent's judgment and into code at the call boundary. The producer no longer approves its own action; policy is enforced where the call crosses. It constrains `criteria_drift` by anchoring policy in code that survives prompt rewrites, persona changes, summarization, and context compaction. The move is policy at the boundary, not in the prose. ## Observable Signal Every Guardrail Decorator report should include: * call site, such as model identity, tool name, or retriever name; * registered policies in precedence order; * policies checked before the decision; * hook fired (`before`, `after`, `on_error`, or `none`); * policy that fired, or `none`; * decision (`pass`, `deny`, `replace`, `sanitize`, `recover`); * args; * original_response, or `not_invoked` if a before-hook denied; * override_response, or `none` for pass-through. A useful report names the blocked call: ```text call_site: tool:delete_file policies: [deny_protected_path, scope_guard, audit_logger] checked: [deny_protected_path] hook_fired: before policy: deny_protected_path decision: deny args: {"path": "/etc/passwd"} original_response: not_invoked override_response: {"error": "protected path denied by boundary policy"} ``` ## Failure Modes * **Prompt-Only Policy:** the policy text lives in the system prompt, and model compliance is the only enforcement. Move the rule into a registered callback at the call boundary. * **No-Op Hook:** a decorator exists but always returns `None`. The interface is policy-shaped, the implementation is pass-through. Add a positive deny-path test and assert the original call was not invoked. * **Hook Without Decision Log:** the wrapper short-circuits silently. Stamp which policy fired, which hook fired, and what override was returned. * **Adapter Mistaken For Guardrail:** a type-conversion wrapper is labeled "guardrail" but does not enforce policy. Type adaptation belongs on Tool Adapter. Rename the wrapper to match its job. ## Use When Use this pattern when: * the framework supports lifecycle hooks at model, tool, retriever, or output boundaries; * policy enforcement gates side effects such as file writes, network calls, deletions, charges, sends, or deploys; * policy needs to survive prompt rewrites, persona tests, and context compaction; * audit requires a decision log per call; * the policy can be expressed as a deterministic decision function rather than a subjective judgment. ## Do Not Use When Do not reach for Guardrail Decorator when: * the policy is genuinely subjective and a **Judge Harness** is the right verifier; * the framework's call boundaries cannot be wrapped; * a one-off conditional in the call site is clearer than registering a decorator; * shape normalization is the actual need, which belongs on Tool Adapter. If hook surfaces are unavailable, label the policy as advisory and add a separate executable check downstream. Do not imply enforcement that does not exist. ## Evidence * **Verification Design Principle 1:** the design doc names external signals over self-review. Prompt-only policy has the same failure shape: the model judges the policy it is asked to obey. * **Verification Design Principle 6:** the design doc treats executable checks as the strongest verification move. A registered callback is the policy-side analog: executable enforcement instead of verbal instruction. * **[ADK](https://github.com/google/adk-python) plugin model callbacks:** the guardrail causal sweep records model-boundary callbacks with before, after, and error hooks. Non-`None` returns halt later callbacks and propagate upward. * **ADK plugin tool callbacks:** the same sweep records tool-boundary callbacks where a before-tool callback can skip the tool call and an after-tool callback can replace the result. * **No-credible-antipattern result:** the antipattern cleanup sweep inspected ADK plugin examples and [AutoGPT](https://github.com/Significant-Gravitas/AutoGPT) `validate_url` without finding a no-op callback instance to promote. ## Related Patterns * **Constitution:** defines the rubric the guardrail enforces. Criteria source and boundary enforcement compose. * **Tool Adapter:** normalizes call shape. Guardrail Decorator enforces call policy. Same wrapper surface, different job. * **Causal Tag:** callback context carries IDs that make the decision log queryable. The guardrail decision becomes another stamped event on the trace. * **Trajectory Cursor:** cursor and policy hooks share the same boundary. The cursor advances after the policy permits, not before. * **Admissibility Gate:** post-call sanitization can execute default-no rejection rules on tool or model outputs. --- Canonical: https://verificationdesign.com/patterns/context-and-state/causal-tag/ Source: ai-design-patterns/cards/causal-tag.md # Causal Tag *(Context Pattern)* ## Name **Causal Tag** Also known as: Causal ID, Run ID Propagation, Invocation Tag, Event Parentage. ## Intent Stamp emitted events with stable joinable identifiers and parent links, then join destination observations to recorded actions; causal attribution depends on trustworthy tag generation and propagation. ## Problem Agents often act into shared event surfaces: logs, traces, message buses, third-party APIs, file stores, and callback systems. A verifier looking at that surface has to answer a causality question: Did this agent cause the observed effect, or did something similar happen nearby? Without causal tags, attribution usually falls back to a time window: * capture `start_time`; * trigger the agent; * count matching log entries after `start_time`; * pass if the count is high enough. That is weaker than it looks. Any producer that emits a matching event after the timestamp can satisfy the assertion. The verifier has a time baseline, but no identity. This is the same ambient-state problem as **State Baseline** and **Delta**, but the failure surface is distributed and asynchronous. `verification_design.md` Principle 9 says assertions must prove the system's action caused the expected outcome, not that the environment happened to contain matching data. Causal Tag is the identity-based answer: make the agent's own effects queryable. ## Forces * **Framework tracing vs. custom tagging.** Many frameworks already expose run IDs, invocation IDs, or callback parentage. Bypassing that layer creates unnecessary attribution debt. * **Time-window attribution vs. identity attribution.** Timestamps narrow the search space; fresh, run-scoped IDs support attribution when the verifier requires that ID on the observed effect. * **Flat ID vs. tree.** A single run ID groups events. A parent ID reconstructs causality across nested model, tool, retriever, and guardrail calls. * **Generated IDs vs. supplied IDs.** Framework-generated IDs reduce caller burden; caller-supplied IDs help join to external payloads. * **Propagation cost vs. audit cost.** Carrying IDs through each boundary adds plumbing, but missing IDs make failures expensive to investigate. ## Solution Tag every event the agent emits with a stable identifier. When events form a causal chain, record the parent identifier too. Use the framework tracing layer if one exists; do not invent a parallel ID system unless the framework has none. The pattern lives at three layers: * **Framework callback signature:** model, tool, and retriever callbacks receive `run_id` and optional `parent_run_id`. LangChain's callback tree is the canonical shape. * **Invocation-scoped context:** a context object owns a stable invocation ID; session events inherit it; each event also gets its own ID. ADK's `InvocationContext` and `Event.id` are the canonical variant. * **Domain payload:** outbound side effects carry the tag into the external surface: message attributes, HTTP headers, log fields, third-party metadata, or filenames. Framework integration is half the pattern. The other half is propagation into the surface the verifier actually reads. A `run_id` trapped inside an in-memory trace store does not help a verifier filtering Cloud Run logs unless the same ID appears in those logs. ## Mechanism 1. **Identify the tracing layer.** Use callback tree, invocation context, or event store. If none exists, create a small context object with a stable ID at run entry. 2. **Use IDs at every boundary.** Model calls, tool calls, retriever calls, guardrail interventions, and outbound side effects all carry the run or invocation ID. 3. **Preserve parent links.** Child events record the parent run or event ID so the tree can be reconstructed. 4. **Stamp emitted artifacts.** Logs, message payloads, API calls, and trace records include the tag in a queryable field. 5. **Verify the tree and destination join.** The report includes ID generation policy, tag propagation, and parent consistency, then joins tags read from the destination surface back to the source action. Orphan events are failures. ## Pattern / Antipattern The same task: verify that a triggered event reached a downstream handler. The antipattern filters shared logs by path and timestamp. The pattern tags the causal chain and verifies events by identity. ### Antipattern: path plus timestamp count The naive implementation captures a time baseline and counts matching log entries after that time. It has partial Delta machinery, but no causal identity. ```python from datetime import datetime, timezone def count_matching_log_entries(logs, path: str, start_time: datetime) -> int: return sum( 1 for entry in logs if entry.path == path and entry.status == 200 and entry.timestamp >= start_time ) def test_trigger_reaches_handler(pubsub, logs): start_time = datetime.now(timezone.utc) pubsub.publish("hello-pipeline-test") count = count_matching_log_entries( logs, path="/apps/trigger_echo_agent/trigger/pubsub", start_time=start_time, ) assert count >= 1 ``` Any matching request after `start_time` satisfies this test. The event may have come from this publish, another test, a retry, a manual request, or ambient traffic. ADK's remote trigger tests have this shape: they publish Pub/Sub or Eventarc messages, then count Cloud Run log entries filtered by request path, status, and timestamp. The test captures a time baseline, so this is not a total absence of verification structure. The missing piece is identity. A unique test-run ID or message ID should be carried in the payload and required in the observed log entry. ### Pattern: run ID tree with parent links The structured implementation records every event with a run ID and parent run ID, carries the tag into the shared destination log, then verifies the destination entry by identity. ```python from dataclasses import dataclass from uuid import UUID, uuid4 @dataclass(frozen=True) class TracedEvent: run_id: UUID parent_run_id: UUID | None kind: str payload: dict class Tracer: def __init__(self): self.events: list[TracedEvent] = [] def start( self, kind: str, parent_run_id: UUID | None = None, run_id: UUID | None = None, **payload, ) -> UUID: run_id = run_id or uuid4() self.events.append( TracedEvent( run_id=run_id, parent_run_id=parent_run_id, kind=f"{kind}.start", payload=payload, ) ) return run_id def emit(self, run_id: UUID, kind: str, **payload) -> None: parent = self.parent_of(run_id) self.events.append( TracedEvent( run_id=run_id, parent_run_id=parent, kind=kind, payload=payload, ) ) def parent_of(self, run_id: UUID) -> UUID | None: for event in self.events: if event.run_id == run_id: return event.parent_run_id raise KeyError(f"Unknown run id: {run_id}") def path_to_root(self, run_id: UUID) -> list[UUID]: path = [] visited = set() while run_id is not None: if run_id in visited: raise ValueError(f"parent cycle at {run_id}") visited.add(run_id) path.append(run_id) run_id = self.parent_of(run_id) return path def check_tree(self) -> dict[str, int | bool]: known_ids = {event.run_id for event in self.events} parents_by_run: dict[UUID, set[UUID | None]] = {} orphan_events = 0 root_parent_errors = 0 for event in self.events: parents_by_run.setdefault(event.run_id, set()).add(event.parent_run_id) if event.parent_run_id is not None and event.parent_run_id not in known_ids: orphan_events += 1 if event.kind == "agent.start" and event.parent_run_id is not None: root_parent_errors += 1 parent_conflicts = sum( 1 for parents in parents_by_run.values() if len(parents) > 1 ) cycles = set() for run_id in known_ids: # Follow each recorded parent, including conflicting parent links. pending = [(run_id, [], set())] while pending: current, path, visited = pending.pop() if current in visited: cycle = path[path.index(current):] first = min(range(len(cycle)), key=lambda i: cycle[i].int) cycles.add(tuple(cycle[first:] + cycle[:first])) continue for parent in parents_by_run.get(current, set()): if parent is not None: pending.append((parent, path + [current], visited | {current})) parent_cycles = len(cycles) return { "consistent": ( orphan_events == 0 and parent_conflicts == 0 and root_parent_errors == 0 and parent_cycles == 0 ), "orphan_events": orphan_events, "parent_conflicts": parent_conflicts, "root_parent_errors": root_parent_errors, "parent_cycles": parent_cycles, } def write_handler_log( tracer: Tracer, destination_log: list[dict], run_id: UUID, path: str, *, propagate_tag: bool, ) -> dict: payload = {"source": "agent"} if propagate_tag: payload["causal_tag"] = str(run_id) else: payload["expected_causal_tag"] = str(run_id) entry = { "path": path, "status": 200, "payload": payload, } destination_log.append(entry) if propagate_tag: tracer.emit( run_id, "handler.received", surface="shared_log", causal_tag=str(run_id), ) return entry def count_by_path_and_status(destination_log: list[dict], path: str) -> int: return sum( 1 for entry in destination_log if entry["path"] == path and entry["status"] == 200 ) def query_by_causal_tag(destination_log: list[dict], run_id: UUID) -> list[dict]: return [ entry for entry in destination_log if entry["payload"].get("causal_tag") == str(run_id) ] def count_missing_destination_tags(destination_log: list[dict]) -> int: return sum( 1 for entry in destination_log if "expected_causal_tag" in entry["payload"] and "causal_tag" not in entry["payload"] ) def short(run_id: UUID | None) -> str: return "null" if run_id is None else run_id.hex[-4:] def render_report( tracer: Tracer, root_id: UUID, *, destination_tag_missing: int, id_generation_policy: str, ) -> str: tree = tracer.check_tree() lines = [f"root_run_id: {short(root_id)}"] for event in tracer.events: surfaces = ["trace_store"] if event.payload.get("causal_tag") == str(event.run_id): surfaces.append("shared_log_payload") lines.append( "event: " f"{event.kind} run: {short(event.run_id)} " f"parent: {short(event.parent_run_id)} " f"tag_present: {', '.join(surfaces)}" ) lines.extend( [ "propagation_surfaces_checked: trace_store, shared_log_payload", "destination_tag_observed: shared_log_payload", f"tree_consistency: {'pass' if tree['consistent'] else 'fail'}", f"orphan_events: {tree['orphan_events']}", f"parent_cycles: {tree['parent_cycles']}", f"destination_tag_missing: {destination_tag_missing}", f"id_generation_policy: {id_generation_policy}", ] ) return "\n".join(lines) tracer = Tracer() root_id = tracer.start( "agent", run_id=UUID("00000000-0000-0000-0000-00000000a001"), prompt="answer user", ) tool_id = tracer.start( "tool", parent_run_id=root_id, run_id=UUID("00000000-0000-0000-0000-00000000b002"), name="web_fetch", ) path = "/apps/trigger_echo_agent/trigger/pubsub" destination_log = [ { "path": path, "status": 200, "payload": {"source": "ambient"}, } ] agent_entry = write_handler_log( tracer, destination_log, tool_id, path, propagate_tag=True ) dropped_tag_entry = write_handler_log( tracer, destination_log, tool_id, path, propagate_tag=False ) selected = query_by_causal_tag(destination_log, tool_id) assert count_by_path_and_status(destination_log, path) > 1 assert selected == [agent_entry] assert agent_entry["payload"]["causal_tag"] == str(tool_id) assert "causal_tag" not in destination_log[0]["payload"] assert dropped_tag_entry not in selected assert count_missing_destination_tags(destination_log) == 1 assert tracer.check_tree() == { "consistent": True, "orphan_events": 0, "parent_conflicts": 0, "root_parent_errors": 0, "parent_cycles": 0, } events_with_orphan = list(tracer.events) events_with_orphan.append( TracedEvent( run_id=UUID("00000000-0000-0000-0000-00000000c003"), parent_run_id=UUID("00000000-0000-0000-0000-00000000d004"), kind="tool.start", payload={}, ) ) orphan_demo = Tracer() orphan_demo.events = events_with_orphan assert orphan_demo.check_tree()["consistent"] is False assert orphan_demo.check_tree()["orphan_events"] == 1 cycle_demo = Tracer() a, b = uuid4(), uuid4() cycle_demo.start("tool", run_id=a, parent_run_id=b) cycle_demo.start("tool", run_id=b, parent_run_id=a) assert cycle_demo.check_tree()["consistent"] is False assert cycle_demo.check_tree()["parent_cycles"] == 1 try: cycle_demo.path_to_root(a) except ValueError as error: assert str(error) == f"parent cycle at {a}" else: raise AssertionError("cyclic ancestry must fail") expected_report = """root_run_id: a001 event: agent.start run: a001 parent: null tag_present: trace_store event: tool.start run: b002 parent: a001 tag_present: trace_store event: handler.received run: b002 parent: a001 tag_present: trace_store, shared_log_payload propagation_surfaces_checked: trace_store, shared_log_payload destination_tag_observed: shared_log_payload tree_consistency: pass orphan_events: 0 parent_cycles: 0 destination_tag_missing: 1 id_generation_policy: caller_supplied""" assert render_report( tracer, root_id, destination_tag_missing=count_missing_destination_tags(destination_log), id_generation_policy="caller_supplied", ) == expected_report ``` LangChain's tracing APIs demonstrate this at framework level: chat model, tool, and retriever callbacks receive `run_id`, optional `parent_run_id`, tags, and metadata. Event-stream tracing stores run and parent maps, emits events with run IDs and parent IDs, and assigns a UUID when no run ID is supplied. ADK uses the invocation-scoped variant. `InvocationContext` owns an `invocation_id`, session events can be filtered to the current invocation, each `Event` has an `invocation_id`, and each event has its own unique `id`. Different shape, same purpose: make causality queryable. ## Determinism Move Causal Tag constrains `ambient_state` by stamping every event with a queryable identity. The verifier filters the agent's contributions out of shared logs, traces, and message buses instead of accepting any matching event. It constrains `async_timing` by replacing "this happened around the same time" with "this event carries the parent ID of the action." Parentage remains queryable across batching, reordering, delayed delivery, and concurrent producers. ## Observable Signal Every Causal Tag report should include: * run or invocation ID; * parent run or event ID, null only for roots; * propagation surfaces checked, such as trace store, log payload, message attribute, API header; * tag presence per surface; * tree consistency result; * orphan-event count, parent-cycle count, and destination-tag-miss count; * ID generation policy (`caller_supplied`, `framework_assigned`, `uuid_on_missing`). A useful report is a small join table: ```text root_run_id: a001 event: agent.start run: a001 parent: null tag_present: trace_store event: tool.start run: b002 parent: a001 tag_present: trace_store event: handler.received run: b002 parent: a001 tag_present: trace_store, shared_log_payload propagation_surfaces_checked: trace_store, shared_log_payload destination_tag_observed: shared_log_payload tree_consistency: pass orphan_events: 0 parent_cycles: 0 destination_tag_missing: 1 id_generation_policy: caller_supplied ``` ## Failure Modes * **Cyclic Parent Links:** An event's ancestry loops back on itself, usually from retries or handlers that reuse a run ID as their own parent. A tree check that only tests orphans and conflicts reports the loop as consistent, and any root walk hangs. Detect revisits and fail the tree. * **Time-Window Attribution:** Tests filter shared logs by path and timestamp. Any matching event after the start time can satisfy the assertion. Include a unique tag in the payload and require the observed event to carry it. * **Framework Bypass:** The framework offers `run_id` and `parent_run_id`, but custom code emits events outside that path. Use the framework callback or explicitly propagate the IDs. * **Flat ID Tree:** Every event has an ID, but no parent link. The verifier sees a bag of events, not a causal chain. Record parent IDs and assert tree consistency. * **Tagged at Source, Untagged at Destination:** The outbound message has an ID, but the downstream log or callback drops it. Verify propagation across the boundary, not only at the source. ## Use When Use this pattern when: * the agent emits events into shared logs, traces, message buses, APIs, or side-effect targets; * multiple agents, runs, or producers share the same event surface; * async timing, batching, or reordering makes temporal proximity unreliable; * verification must attribute an observed effect to a specific action; * a framework tracing layer exists but is not yet engaged. ## Do Not Use When Do not reach for Causal Tag when: * the event surface is fully private to the test or run; * **State Baseline** plus **Delta** gives equivalent attribution more cheaply; * the boundary cannot carry any metadata and cannot be changed; * the side effect is low stakes and misattribution cost is negligible. When tags cannot propagate, redesign the boundary or fall back to snapshot-based attribution. ## Evidence * **Verification Design Principle 9:** The design doc requires assertions to prove that the system's action caused the outcome. Causal Tag is the identity-based partner to State Baseline's snapshot-based answer. * **[LangChain](https://github.com/langchain-ai/langchain) tracing:** The evidence summary records callback APIs that carry `run_id`, `parent_run_id`, tags, and metadata through chat model, tool, and retriever events. * **LangChain event stream:** The same evidence records run maps, parent maps, emitted run IDs, parent IDs, and UUID assignment when run IDs are missing. * **[ADK](https://github.com/google/adk-python) InvocationContext and Event:** The evidence summary records invocation IDs, per-event IDs, current-invocation event filtering, and function-response resolution by invocation. * **ADK remote trigger tests:** The Delta sweep records the antipattern: Cloud Run logs are counted by path and timestamp without a unique message or test-run ID. ## Related Patterns * **Delta:** requires tagged events to attribute changes. Delta on untagged shared logs is the Causal Tag antipattern. * **State Baseline:** provides snapshot-based attribution when tag propagation is infeasible. * **Trajectory Cursor:** cursor entries should carry causal tags so distributed events can be joined back to steps. * **Guardrail Decorator:** guardrail interventions should emit tagged parent events. * **Executable Analog:** tag-presence and tree-consistency checks are executable checks over the trace surface. --- Canonical: https://verificationdesign.com/patterns/context-and-state/trajectory-cursor/ Source: ai-design-patterns/cards/trajectory-cursor.md # Trajectory Cursor *(Context Pattern)* ## Name **Trajectory Cursor** Also known as: Action History Cursor, Step Cursor, Trajectory Record, Execution Cursor. ## Intent Maintain an explicit, structured record of where the agent is in its multi-step process and what happened at each boundary, so the verifier and the next turn can read the trajectory instead of inferring it from chat history or model recall. ## Problem Agents do not fail only by choosing a bad action. They also fail by losing track of what happened. In a multi-step loop, the model may propose an action, a policy layer may deny it, a tool may error, a retry may run, or a fallback may skip the intended path. If that boundary is not recorded in the trajectory, the next turn reconstructs the past from chat text and latent memory. The missing event becomes invisible. The agent may propose the same denied command again, skip the action that mattered, or claim progress that never occurred. The verifier has the same problem. A final answer may look plausible, but without a cursor the verifier cannot tell which tool calls actually happened, which errors were recovered, which denials were seen by the next turn, or whether the loop made forward progress. `verification_design.md` Principle 3 frames tool-agent verification as a step-level problem. Its ToolPRMBench update points toward checking tool choice, arguments, and observed tool-state transitions at each action boundary, using interaction history rather than final task success alone (arXiv:2601.12294). Trajectory Cursor is the context pattern that makes that interaction history explicit. ## Forces * **Model recall vs. explicit record.** Chat history contains the story, but a cursor records the state machine the verifier can inspect. * **Cursor granularity.** Some systems need one entry per tool call; others need per-step, per-task, or per-graph-node state. * **Forward-only vs. resumable.** A short loop may only need append-only history; a workflow engine may need persisted cursor state for pause and resume. * **Audit log vs. control signal.** A passive transcript helps later review, but a cursor should also guide the next action. * **Cursor maintenance cost vs. loop failure cost.** Recording each boundary adds bookkeeping; missing one denial can create an infinite retry loop. ## Solution Define a structured trajectory schema and append one entry at every observable boundary: proposal, decision, execution, denial, result, and feedback. Before the agent composes its next action, read the cursor as the source of truth. Do not ask the model to infer what happened from prose. Do not let failure paths return bare errors that never enter history. If a permission check denies an action, the denial is an entry. If a retry happens, the retry is an entry. If a tool result is missing, the missing result is visible as a gap. Common shapes: * **Action-history list:** entries record tool name, arguments, decision, outcome, error, denial reason, and feedback. This is the canonical shape for single-agent loops. * **Step counter with status:** planned, executed, denied, skipped, errored. This is enough when the action graph is fixed. * **Resumable snapshot:** persisted cursor state lets a workflow pause, replay a prefix, and resume. Dify's workflow event snapshot evidence fits this variant. * **Graph execution cursor:** per-task or per-node execution state separates reusable agent objects from the graph's current execution position. AutoGen GraphFlow tests are the canonical pattern evidence here. The cursor is not just a log. A log can be written after the fact. A cursor is consumed by the next turn and by the verifier. ## Mechanism 1. **Define the schema.** Name valid entry types such as proposal, decision, execution, denial, result, and feedback. 2. **Append before yielding.** Every boundary appends a structured entry before control returns to the next turn or node. 3. **Read before proposing.** The next action is conditioned on the recorded trajectory, not on model recall. 4. **Verify forward progress.** Detect repeated proposals, missing critical actions, unrecorded boundaries, and stalled cursor positions. 5. **Report cursor state.** Record the last cursor position, the entry that advanced it, and any progress checks the verifier ran. ## Pattern / Antipattern The two examples show the same cursor move at different boundary types: the antipattern loses a denial outside the trajectory, while the pattern records each graph-node boundary on the observed outcome that advanced it. ### Antipattern: denial outside the trajectory The naive implementation checks permission, receives a denial, returns a bare error, and leaves action history unchanged. The next turn cannot see the denial and may propose the same command again. ```python class AgentLoop: def __init__(self, model, permission_manager): self.model = model self.permission_manager = permission_manager self.action_history = [] def run_turn(self, task): proposal = self.model.propose(task, history=self.action_history) decision = self.permission_manager.check(proposal) if not decision.allowed: return { "status": "error", "message": "Permission denied", } result = proposal.execute() self.action_history.append({ "type": "result", "tool": proposal.tool_name, "args": proposal.args, "result": result, }) return result ``` The denial is real, but the cursor never sees it. The model receives the same `action_history` on the next turn and can re-propose the same denied action. AutoGPT's permission-denial regression test names this exact failure: a denied command returned an error without feedback in action history, so the agent had no memory of the denial and proposed the same command again. The fix was to register the denial as feedback through `do_not_execute`, making the boundary visible to the next prompt. AutoGPT's benchmark analyzer shows the same class at population scale: repeated same-tool loops, missing critical actions such as `write_file`, timeouts, and unrecovered errors are trajectory-level failures, not merely bad final answers. ### Pattern: cursor advances at every boundary The structured implementation records the graph execution cursor and verifies that each task advances through the expected path. ```python class EchoAgent: def __init__(self, name): self.name = name def run(self, task): return f"{self.name}:{task}" class GraphCursor: def __init__(self, nodes): self.nodes = nodes self.entries = [] def run_task(self, task): task_id = len({entry["task_id"] for entry in self.entries}) + 1 self.entries.append({ "type": "proposal", "task_id": task_id, "content": task, }) for node in self.nodes: outcome = node.run(task) self.entries.append({ "type": "execution", "task_id": task_id, "node": node.name, "outcome": outcome, }) return [entry for entry in self.entries if entry["task_id"] == task_id] def report(self, task_id): entries = [entry for entry in self.entries if entry["task_id"] == task_id] executions = [entry for entry in entries if entry["type"] == "execution"] expected_nodes = [node.name for node in self.nodes] recorded_nodes = [entry["node"] for entry in executions] missing_nodes = [node for node in expected_nodes if node not in recorded_nodes] repeat_count = sum( 1 for previous, current in zip(recorded_nodes, recorded_nodes[1:]) if previous == current ) last_entry = executions[-1] return { "cursor": f"graph.task.{task_id}", "last_entry": f"execution(node={last_entry['node']}, task_id={task_id})", "path": " -> ".join(recorded_nodes), "repeat_loop_count": repeat_count, "missing_required_actions": missing_nodes, "unrecorded_boundaries": len(expected_nodes) - len(executions), } agents = [EchoAgent("A"), EchoAgent("B"), EchoAgent("C")] cursor = GraphCursor(nodes=agents) first = cursor.run_task("First task") second = cursor.run_task("Second task") assert [entry["node"] for entry in first if entry["type"] == "execution"] == ["A", "B", "C"] assert [entry["outcome"] for entry in second if entry["type"] == "execution"] == [ "A:Second task", "B:Second task", "C:Second task", ] assert {entry["task_id"] for entry in first}.isdisjoint( {entry["task_id"] for entry in second} ) assert cursor.report(2) == { "cursor": "graph.task.2", "last_entry": "execution(node=C, task_id=2)", "path": "A -> B -> C", "repeat_loop_count": 0, "missing_required_actions": [], "unrecorded_boundaries": 0, } ``` AutoGen's GraphFlow test has this shape: the same graph runs two sequential tasks, each result contains the expected user, A, B, C message path, and each agent's message counter increments once per task. The cursor is the graph's per-task execution position, not the agent objects themselves. Dify's workflow event snapshot service is a resumable variant: persisted workflow state can reconstruct an event prefix before live execution continues. AutoGen's C# termination tests are supporting evidence for the reset shape, but the Python GraphFlow test is the canonical card instance. ## Determinism Move Trajectory Cursor constrains `tool_boundary_ambiguity` by requiring every observable boundary to produce a structured entry. The next turn reads the explicit trajectory record instead of reconstructing proposals, denials, executions, and results from chat memory. The determinism move is boundary accounting. A proposal, denial, execution, result, retry, and feedback event are different states. If the cursor records them, the next turn and verifier can distinguish them. If not, the model is left to guess what happened. ## Observable Signal Every Trajectory Cursor report should include: * `cursor`: the last cursor position; * `last_entry`: the entry that advanced the cursor most recently; * `path`: the recorded path for the task; * `repeat_loop_count`: repeated adjacent entries in the recorded path; * `missing_required_actions`: expected actions absent from the recorded path; * `unrecorded_boundaries`: expected boundaries with no execution entry. The most useful report shows forward progress directly: ```text cursor: graph.task.2 last_entry: execution(node=C, task_id=2) path: A -> B -> C repeat_loop_count: 0 missing_required_actions: [] unrecorded_boundaries: 0 ``` ## Failure Modes * **Repeat-Loop:** The agent repeats an action because the previous success, denial, or error is not visible to the next turn. Append the outcome before dispatching the next turn. * **Missing Critical Action:** The cursor advances without invoking the action required to satisfy the task. Track required actions and assert that each appears in the trajectory. * **Unrecorded Boundary:** Permission denials, retries, fallbacks, or tool errors occur outside the cursor. Require every boundary, including failure paths, to emit an entry. * **Cursor Drift:** The cursor advances because the model says progress happened, but no execution entry records the completed step. Advance only on observable outcomes. ## Use When Use this pattern when: * an agent runs more than two tool calls per task; * permission, retry, fallback, or denial paths exist outside the main action flow; * verification needs to prove which steps executed, not which the agent narrated; * a loop can stall, repeat, or skip required actions; * pause, replay, or human-in-the-loop resume matters. ## Do Not Use When Do not reach for Trajectory Cursor when: * the workflow is a single-shot model call with no loop; * the framework already persists a structured trajectory and another cursor would duplicate state; * the process is a deterministic transformation pipeline with no agent decisions; * the only verification question is final-state causality, where **State Baseline** or **Delta** is enough. ## Evidence * **Verification Design Principle 3:** The design doc's ToolPRMBench update frames tool-agent verification around interaction history and action-boundary checks, not only final success (arXiv:2601.12294). * **[AutoGen](https://github.com/microsoft/autogen) GraphFlow:** The evidence summary records a Python GraphFlow test where the same graph runs two sequential tasks, each task follows the expected A, B, C path, and agent counters show one execution per task. * **[Dify](https://github.com/langgenius/dify) workflow snapshots:** The evidence summary records a resumable cursor variant where workflow events are reconstructed from persisted snapshots. * **[AutoGPT](https://github.com/Significant-Gravitas/AutoGPT) permission-denial regression:** The antipattern cleanup sweep records a denied action that was not registered in history, causing repeat proposals until the denial was added as feedback. * **AutoGPT benchmark analyzer:** The same sweep records trajectory-level failure categories: repeated same-tool loops, missing critical actions, timeouts, and unrecovered errors. ## Related Patterns * **State Baseline:** snapshots environment state; Trajectory Cursor snapshots the agent's action history. * **Delta:** forward progress is a Delta on cursor position. * **Causal Tag:** links cursor entries to external artifacts the agent created. * **Executable Analog:** repeat-loop and missing-action checks are executable checks over the cursor. * **Constitution:** can encode required actions and cursor schema as verification criteria. --- Canonical: https://verificationdesign.com/patterns/context-and-state/state-baseline/ Source: ai-design-patterns/cards/state-baseline.md # State Baseline *(Context Pattern)* ## Name **State Baseline** Also known as: Pre-Action Snapshot, Baseline Capture, Snapshot-Restore Harness. ## Intent Capture the relevant environment or process state before an action under verification, so the verifier can establish that observed state did not already exist before the action; causal attribution requires isolation or a Causal Tag. ## Problem Agentic systems act inside shared state: filesystems, global registries, environment variables, databases, queues, browser sessions, and cloud accounts. A verifier that only looks at post-action state cannot tell whether the agent caused the state or merely found it. For example, if an agent is supposed to delete a file, a post-action check like `assert not path.exists()` passes when the file was already absent. If an agent is supposed to register a hook, `assert len(hooks) == 1` passes when a previous test leaked one hook into the registry. The check is green, but it has not verified the action. This is ordinary state contamination, but LLM agents make it easier to miss. A deterministic test usually fails later when inherited state conflicts with another assertion. An agent can rationalize the coincidental pass and continue the run as if it proved success. `verification_design.md` Principle 9 names this explicitly: assertions must prove the system's action caused the expected outcome, not that the environment happened to contain matching data. Luo et al. identify state leakage and order dependence as recurring sources of flaky tests in software systems (FSE 2014). Without a baseline, verification cannot separate causality from coincidence. ## Forces * **Shared environments vs. fresh environments.** A fresh container or mocked database avoids inherited state, but many useful checks run against persistent workspaces or live services. * **Snapshot scope vs. snapshot cost.** Capturing every file, row, or variable can be expensive; capturing too little leaves hidden state surfaces. * **Per-action vs. per-session baselines.** A session-level snapshot can support rollback, but a verifier often needs a baseline immediately before a specific action. * **Comparison vs. cleanup.** Some baselines exist to compute a Delta; others exist to restore state after the action. * **Local mutation vs. audit trail.** Direct mutation is simple, but destructive operations need a preimage if they are part of a verification loop. ## Solution Snapshot the relevant state surface before the agent acts. After the action, either compare post-state to the baseline or restore the baseline during cleanup. The baseline must be close enough to the action to answer the verifier's actual question: "what changed during this step?" A startup snapshot may be useful for session recovery, but it does not prove a particular action caused a particular post-condition. Common shapes: * **In-memory snapshot:** copy global lists, registries, caches, or settings before the test, then restore them after. * **File-backed inventory:** write a named list of files, resources, plugin IDs, or API objects, then compare the current inventory to that baseline. * **Process-environment snapshot:** capture environment variables and redirect state directories to a temporary location, then restore in `finally`. * **Snapshot-replay:** persist enough runtime state to reconstruct an event stream or resume from a pause boundary. * **Git preimage:** commit or hash original files before an automated edit, then compare or restore from that preimage. State Baseline is often paired with **Delta**. State Baseline supplies the pre-state. Delta computes the change. ## Mechanism 1. **Identify the state surface.** Name the files, registry, environment variables, database rows, queue messages, or runtime objects that could satisfy the assertion by accident. 2. **Capture pre-state.** Snapshot that state immediately before the action under verification, using the narrowest mechanism that covers the risk. 3. **Run the action.** Hand control to the agent, tool, workflow, or test body. 4. **Measure or restore.** Either capture post-state for comparison or restore the pre-state during cleanup. 5. **Report the baseline.** Record the snapshot mechanism, baseline value or reference, post-action value, and cleanup status. ## Pattern / Antipattern Both examples show the same State Baseline move on different state surfaces. The antipattern mutates workspace files with no nearby preimage. The pattern snapshots an in-memory registry, clears inherited state for the action, and restores the original registry afterward. ### Antipattern: mutate shared state with no preimage This shape is not wrong in every production path. It becomes a verification antipattern when an agent-facing operation can alter shared state and the surrounding harness has no nearby baseline, restore point, or diff check. ```python from pathlib import Path class WorkspaceFiles: def __init__(self, root: Path): self.root = root def move(self, source: str, destination: str) -> None: source_path = self.root / source destination_path = self.root / destination destination_path.parent.mkdir(parents=True, exist_ok=True) source_path.rename(destination_path) def delete(self, path: str) -> None: target = self.root / path target.unlink() ``` If a verifier later checks only `not source_path.exists()` or `destination_path.exists()`, it cannot prove this operation caused that post-state. The file may have already been absent, already moved, or modified by another actor. The verifier has no preimage to inspect. AutoGPT's file manager evidence has this shape: append, move, and delete mutate workspace files directly, while `save_state` exists as an explicit session operation rather than a per-action baseline. Treat it as a partial-fit antipattern because direct mutation is legitimate outside a verification seam; the risk appears when an agent loop treats post-state as proof without a nearby snapshot. ### Pattern: snapshot, isolate, restore The structured implementation captures the original registry, clears inherited state, yields to the action, and restores the exact pre-state afterward. ```python from contextlib import contextmanager before_hooks = [] after_hooks = [] error_hooks = [] def hook_counts(): return [len(before_hooks), len(after_hooks), len(error_hooks)] def report(pre_state, post_state): return { "check": "hooks.registry.isolated", "snapshot": "in_memory", "covered": ["before_hooks", "after_hooks", "error_hooks"], "pre_state": pre_state, "post_state": post_state, "cleanup": "restored", } @contextmanager def isolated_hooks(): # 1. Capture the pre-state. baseline = (list(before_hooks), list(after_hooks), list(error_hooks)) # 2. Remove inherited state for this test. before_hooks.clear() after_hooks.clear() error_hooks.clear() try: # 3. Run the action under verification. yield finally: # 4. Restore the pre-state even if the test fails. before_hooks.clear() after_hooks.clear() error_hooks.clear() before_hooks.extend(baseline[0]) after_hooks.extend(baseline[1]) error_hooks.extend(baseline[2]) before_hooks.append("inherited") captured_baseline = (list(before_hooks), list(after_hooks), list(error_hooks)) with isolated_hooks(): after_hooks.append("registered") isolated_state = hook_counts() assert "inherited" not in before_hooks and after_hooks == ["registered"] assert (before_hooks, after_hooks, error_hooks) == captured_baseline baseline_report = report( pre_state=[len(hooks) for hooks in captured_baseline], post_state=isolated_state, ) assert baseline_report == { "check": "hooks.registry.isolated", "snapshot": "in_memory", "covered": ["before_hooks", "after_hooks", "error_hooks"], "pre_state": [1, 0, 0], "post_state": [0, 1, 0], "cleanup": "restored", } ``` The verifier can now reason about state inside the test body without inheriting previous runs. CrewAI's hook tests use this in-memory snapshot-and-restore shape around global hook lists. Other sourced variants use a file-backed inventory, a temporary home directory with environment restore, workflow event snapshots, and owned-resource ledgers. ## Determinism Move State Baseline constrains `ambient_state` by recording the state the agent did not create. Restoring the baseline after the action closes the same source of error: one run's residue should not become the next run's setup. The move is simple: make inherited state observable before the agent acts. Once pre-state is explicit, a verifier can compute Delta, restore cleanup state, or reject a post-condition that was already true. ## Observable Signal Every State Baseline report should include: * `check`: the verifier or registry check being protected; * `snapshot`: the snapshot mechanism (`in_memory`, `file_backed`, `process_env`, `snapshot_replay`, `git_preimage`); * `covered`: the state surfaces covered by the baseline; * `pre_state`: the pre-action value or snapshot reference; * `post_state`: the post-action value when comparison is used; * `cleanup`: the cleanup status (`restored`, `leaked`, `partial`, `not_applicable`). The most useful report is boring and concrete: ```text check: hooks.registry.isolated snapshot: in_memory covered: ["before_hooks", "after_hooks", "error_hooks"] pre_state: [1, 0, 0] post_state: [0, 1, 0] cleanup: restored ``` ## Failure Modes * **Stale Baseline:** The snapshot is captured too early, such as at process start, while the action is verified much later. Capture baselines in the narrowest scope possible. * **Snapshot Leak:** The test fails before cleanup and leaves modified state behind. Use `try/finally`, pytest yield fixtures, or equivalent cleanup mechanisms. * **Partial Snapshot:** The harness captures one state surface but misses another. Enumerate coverage in the report and fail when a required surface is uncovered. * **Baseline Without Version:** A stored baseline file drifts silently. Version baseline files and require explicit migration notes when the expected inventory changes. ## Use When Use this pattern when: * the environment is shared, persistent, or reused across runs; * a pre-existing object could satisfy the post-condition; * concurrent activity can create, delete, or mutate matching state; * the action is destructive and needs an audit trail; * the verifier needs to prove causation, not merely observe a final condition. ## Do Not Use When Do not reach for State Baseline when: * the environment is fully ephemeral and recreated for each run; * the state surface is too large and the assertion is low stakes; * the verified property is purely internal to a deterministic function call; * a **Causal Tag** can identify the agent's own artifacts more directly; * a pre-state read would itself change the system under test. ## Evidence * **Verification Design Principle 9:** The canonical design doc states that verification must prove the agent's action caused the outcome, not that ambient state already matched. It cites Luo et al. on state contamination in flaky tests (FSE 2014). * **[CrewAI](https://github.com/crewAIInc/crewAI) hook fixture:** The evidence summary records a direct State Baseline pattern: global hook registries are copied, cleared, yielded to the test, then restored. * **[OpenClaw](https://github.com/openclaw/openclaw) and [Dify](https://github.com/langgenius/dify) variants:** The evidence summary records file-backed inventory, process-environment snapshot, and workflow snapshot-replay variants. * **[AutoGPT](https://github.com/Significant-Gravitas/AutoGPT) file manager:** The antipattern cleanup sweep records direct append, move, and delete operations on workspace files without a nearby per-action preimage. The note treats this as moderate evidence because direct mutation is only the failure when used as a verification seam. ## Related Patterns * **Delta:** uses the baseline to compute what changed during this run. * **Causal Tag:** tags artifacts when snapshotting shared state is too expensive or ambiguous. * **Constitution:** can require baseline capture as part of the criteria contract. * **Trajectory Cursor:** tracks position and state inside a multi-step process. * **Executable Analog:** turns the snapshot or restore check into a runnable verifier. --- Canonical: https://verificationdesign.com/patterns/verification/executable-analog/ Source: ai-design-patterns/cards/executable-analog.md # Executable Analog *(Verification Pattern)* ## Name **Executable Analog** Also known as: Mechanical Extraction, Tool-Mediated Verification, Objective Grounding. ## Intent Translate a subjective, language-based verification step into a deterministic, programmatic execution step that yields a binary pass/fail signal independent of the agent's judgment. ## Problem LLMs are vulnerable to sycophancy and hallucinated reasoning, especially when asked to self-correct or evaluate their own outputs. The most common anti-pattern in agentic design is the instruction: *"Review your work and ensure it is correct."* When an agent is asked to "verify the API is healthy" or "check if the host is visible on the page," it may read a text snapshot (like an accessibility tree or a JSON dump), infer that the desired state is present, and return a passing grade without committing to the observed value. The verification step inherits the biases and blind spots of the generation step. Because naive self-review is unreliable without external feedback, verification steps that rely only on the model's interpretation of text should be treated as weak evidence. ## Forces * **Semantic understanding vs. mechanical checks.** LLMs are useful for interpreting messy language, but they are weak evidence for strict pass/fail claims. Executable checks are narrower, but easier to audit. * **Ease of prompting vs. effort of tooling.** It is much faster to write a prompt asking the model to "check the UI" than to write a headless browser evaluation script. * **Confirmation bias.** Models are biased toward accepting confirmatory framing ("Does this look right?"), especially around their own prior work. ## Solution Where possible, replace LLM judgment calls with runnable checks. Find the executable analog for the claim the agent is trying to verify. Instead of asking the model to read an API response, give it a tool that executes `curl -sf` and assert on the exit code. Instead of passing a DOM snapshot to the model to ask if an element is visible, use `browser_evaluate` with a Javascript query selector and strictly compare the returned string. The strongest verification is grounded in something the system can **execute and observe**, not something the agent merely **reads and opines on**. ## Mechanism 1. **Identify the subjective check.** (e.g., *"Does the webpage show the success message?"*) 2. **Determine the extraction tool.** (e.g., Playwright `page.evaluate()`) 3. **Write the deterministic extractor.** (e.g., `() => document.querySelector('.success-toast')?.textContent`) 4. **Implement Extract-then-Compare.** Force the verification system to log the raw extracted value *before* any comparison happens, creating an auditable trail. ## Pattern / Antipattern The same task: verify that a host name is visible on a rendered web page. The antipattern hands the rendered output to an LLM and asks for a judgment. The pattern executes a deterministic extraction and compares the result to a fixed expected value. ### Antipattern: LLM reads the snapshot and decides The naive implementation captures a snapshot of the rendered page and asks the model whether the expected host is visible. The verifier's pass/fail comes from model interpretation of the snapshot, not from a comparison the system can audit. ```python def verify_host_visible(snapshot: str, expected_host: str, model) -> bool: """Ask the model whether the host appears on the page.""" prompt = ( f"Here is the page snapshot:\n{snapshot}\n\n" f"Is the host {expected_host} visible on the page? Answer YES or NO." ) response = model.complete(prompt) return "YES" in response.upper() ``` There is no extracted observed value, no recorded comparison, no audit trail. The snapshot may contain the host, or contain something that looks like the host, or omit it; the verifier cannot distinguish those cases on review. The same model that wrote the page would happily approve it. ### Pattern: extract mechanically, compare strictly, record both The structured implementation forces a deterministic extraction first, prints the observed value, and only then compares. Failure paths are logged with the same audit fields as success paths so zero-failure reports can be scrutinized. ```python from typing import Callable, Any class BrowserStub: def evaluate(self, js: str) -> str: if "document.querySelector('.host')?.textContent" not in js: raise ValueError(f"unsupported selector query: {js}") return "httpbin.org" browser = BrowserStub() class ExecutableAnalog: """ Wraps an executable check to ensure separation of extraction and judgment. Pass/fail comes from the comparison, not model opinion. """ def __init__(self, check_id: str, extractor: Callable[[], Any], expected: Any): self.check_id = check_id self.extractor = extractor self.expected = expected def verify(self) -> dict: try: # 1. Mechanical extraction (no LLM involved) observed = self.extractor() # 2. Strict comparison passed = (observed == self.expected) # 3. Audit trail generation return { "check_id": self.check_id, "passed": passed, "expected": self.expected, "observed": observed, "error": None, } except Exception as e: return { "check_id": self.check_id, "passed": False, "expected": self.expected, "observed": None, "error": str(e), } def get_host_from_dom(): return browser.evaluate("document.querySelector('.host')?.textContent") checker = ExecutableAnalog( check_id="api.host.visible", extractor=get_host_from_dom, expected="httpbin.org", ) pass_report = checker.verify() mismatch_report = ExecutableAnalog( check_id="api.host.visible", extractor=lambda: "example.com", expected="httpbin.org", ).verify() def failing_extractor(): raise RuntimeError("DOM query failed") error_report = ExecutableAnalog( check_id="api.host.visible", extractor=failing_extractor, expected="httpbin.org", ).verify() assert pass_report["passed"] is True and pass_report["observed"] == pass_report["expected"] assert ( mismatch_report["passed"] is False and mismatch_report["observed"] == "example.com" and mismatch_report["error"] is None ) assert ( error_report["passed"] is False and error_report["observed"] is None and isinstance(error_report["error"], str) and error_report["error"] ) ``` The extractor is a single executable boundary the system can replay. The comparison is a single line the system can diff. The report carries everything a reviewer needs to challenge the verdict without rerunning the test. Anthropic-cookbook's text-to-SQL eval shows the same shape in a literal form: parse the SQL, execute it against SQLite, and assert on the returned rows. Aider's linter loop is the repair-loop version: run a narrow mechanical check, carry line-grounded failures back to the agent, and ask for a bounded fix. ## Determinism Move Executable Analog constrains `self_review_bias` (the same agent that produced the artifact no longer judges whether it satisfies the check) and `judge_subjectivity` (the verdict comes from a deterministic equality on extracted values, not from a model's interpretation of rendered output). By forcing extract-then-compare instead of interpret-and-decide, the system loses the freedom to rationalize a coincidental pass. ## Observable Signal Every report includes: * `check_id`: the named check being run; * `expected`: the value the executable analog is checking against; * `observed`: the raw value returned by the extractor, before judgment; * `passed`: the strict comparison result; * `error`: the exception text when extraction fails, otherwise `None`. A reviewer challenging the verdict reads observed, reads expected, and either agrees the comparison was sound or pins down which side was wrong. Zero-failure reports must still record observed values for every check; a green report with no observed values is not auditable. ```text check_id: api.host.visible expected: httpbin.org observed: httpbin.org passed: true error: null check_id: api.host.visible expected: httpbin.org observed: example.com passed: false error: null ``` ## Failure Modes * **Fixed Sleep over Polling:** If the executable analog relies on asynchronous state, developers often add `time.sleep(5)`. This flakes. Wrap executable analogs in a poll-with-timeout mechanism. * **Brittle Selectors:** The executable extraction logic (like a regex or CSS selector) breaks due to minor, valid changes in the output, causing false negatives. * **Over-delegation:** Letting the same agent write the executable analog on the fly. If the agent writes the test for its own code, it may write a tautological test that passes the bug. ## Use When Use this pattern when: * the claim being verified can be expressed as a deterministic check; * the output has structure (DOM, JSON, exit code, log line) that can be queried; * you can write a test rather than just describe one; * the same check will run repeatedly (regression, CI, multi-agent loops); * the verification trail needs to be auditable. ## Do Not Use When Do not reach for an executable analog when: * the property is genuinely subjective and has no programmatic surface (tone, readability, design taste); * the extractor would be more brittle than the LLM judgment it replaces; * the cost of writing the executable check exceeds the cost of one-off human review; * the claim is exploratory and "good" is not yet defined. When no executable analog exists, route the claim explicitly to a validated judge or human reviewer rather than to an inline LLM judgment. ## Evidence * **Agent-as-a-Judge (ICML 2025):** Evaluators with agency (ability to run code) achieved ~90% agreement with human experts, vs ~70% for LLM-as-Judge. (arXiv:2410.10934) * **LLMs Cannot Self-Correct (ICLR 2024):** The most widely replicated finding is that LLM performance degrades with naive self-correction without external execution feedback. (arXiv:2310.01798) ## Related Patterns * **Comparator:** strict equality is the named comparison step in this example; Comparator is the broader family of operators for deciding whether observed satisfies expected. * **Constitution:** an Executable Analog is how one codified rule becomes a runnable check instead of a prose instruction. * **Blind Oracle:** Blind Oracle protects the expected value from draft contamination; Executable Analog is its executable specialization, the strongest form when the comparison can run as code. * **Judge Harness:** when no executable analog exists, route the claim to a validated judge harness rather than an inline LLM judgment. * **Admissibility Gate:** keeps the challenge independent of the producer, which reduces the risk that the agent writes a tautological check for its own work. --- Canonical: https://verificationdesign.com/patterns/verification/blind-oracle/ Source: ai-design-patterns/cards/blind-oracle.md # Blind Oracle *(Verification Pattern)* ## Name **Blind Oracle** Also known as: Independent Derivation, Extract-Then-Compare, Spec-Only Verification, Draft-Blind Judgment. ## Intent Derive expected evidence from the spec, the question, or independent re-execution without conditioning that derivation on the agent's draft, reasoning trace, or shortcut history. The verifier may inspect the draft afterward to extract observed values. The expected side of the comparison must come from a channel the draft cannot influence. The mechanic is structural independence between the derivation path and the production path, not blindness to the artifact. ## Problem A judge prompt places the draft answer in the verifier's working context before asking for a verdict. The verifier's "is this correct?" collapses into "does this look plausible?" because the verifier has nothing independent to compare against. The same blind spots that produced the answer steer the judgment. Common shapes: * judge prompts show `[Submission]` before asking whether the response meets the criteria; * evaluator prompts pass question, submitted answer, and ground truth into one model call, then parse `CORRECT` or `INCORRECT`; * factored verification questions are answered while the draft is still visible to the verifier; * verifiers "do not see the draft" but receive the writer's reasoning trace, which carries the same anchors as the draft. `verification_design.md` Principle 2 names the rule: if the verifier can see the original output, it copies the same errors. The verification step must not be conditioned on the draft. The Chain-of-Verification factored and revised variant, where verification questions were answered without conditioning on the draft, improved FACTSCORE from 55.9 to 71.4 [arXiv:2309.11495]. Single-agent systems cannot get full independence. The writer and verifier share a context window. Blind Oracle is the mitigation: compute expected before reading the draft, separate extraction from judgment, and use tools where possible. ## Forces * **Independent derivation vs. agent convenience.** Re-deriving expected values from the spec costs another pass; reading the draft does not. * **Blindness to draft vs. inspection of artifact.** The verifier may need the artifact to extract observed values. The artifact must not define expected values. * **Spec-derivable vs. spec-underivable.** The pattern only applies where expected can be computed from a separate path. * **Single-context vs. separate-context.** In a single-agent system, the mitigation is order and source of derivation, not literal context separation. * **Reasoning trace as evidence vs. reasoning trace as draft.** A writer's scratchpad anchors the verifier like the draft does. * **Reference value vs. derived expected.** A pre-supplied reference can substitute for derivation only when the verifier derives expected from it before seeing the draft. ## Solution Compute expected on a derivation path that has no read access to the draft, the writer's reasoning, or the writer's shortcut history. Extract observed from the draft after expected exists. Compare the two with a named operator. The pattern lives at three layers: * **Derivation channel:** the function or call that produces expected takes only the spec, question, or independent reference. Drafts, writer reasoning, and writer summaries are out of scope. * **Extraction channel:** a separate pass extracts observed from the draft. Extraction may read the draft; it does not render the verdict. * **Comparison:** a named **Comparator** operator decides verdict from `expected` and `observed`. No model call that takes the draft as input is invoked at comparison time. Three common shapes: * **Extract-then-compare:** extract raw values, print them, then compare to expected. * **Blind re-derivation:** re-derive from scratch without referencing the prior answer, then compare. * **Tool-mediated extraction:** use parsers, queries, calls, or other tools to extract observed values rather than asking the model to interpret rendered output. ## Mechanism 1. **Identify the derivation channel.** Expected may come from spec parsing, independent execution, a pre-supplied reference, or blind re-derivation by a model call that does not receive the draft. 2. **Compute expected first.** Derive expected before any read access to the draft. The derivation function's signature should not accept the draft as a parameter. 3. **Extract observed from the draft.** A mechanical extractor reads the draft and produces a structured observed value. No verdict is rendered here. 4. **Compare with a named operator.** Exact match, regex, JSON distance, trajectory match, or another named operator decides pass or fail. 5. **Stamp derivation source.** The verdict records `derivation_source: "spec"`, `"reference"`, or `"blind_rederivation"`. ## Pattern / Antipattern The same task: decide whether a submitted answer satisfies a question or criterion. The antipattern puts the submission into the judge prompt before expected is derived. The pattern derives expected first, extracts observed second, and compares them with an explicit operator. ### Antipattern: criteria and submission in one judge prompt The naive implementation fuses extraction, comparison, and judgment in one model call. ```python def grade_submission(model, input_text: str, submission: str, criteria: str) -> str: prompt = f""" [Input] {input_text} [Submission] {submission} [Criteria] {criteria} Does the submission meet the criteria? Return CORRECT or INCORRECT and explain your answer. """ return model.complete(prompt) ``` The submission is in the verifier's context before the verifier has derived expected. The verdict can anchor on the draft and then rationalize from the criteria. LangChain's criteria evaluator and scoring evaluator have this prompt shape: criteria and submission are presented together before the model renders a judgment. The scoring variant may also include a reference answer, but adding `[Reference]` does not make the judge blind when `[Submission]` is still in view. Some evaluations must inspect the artifact being graded. The failure mode is not artifact inspection; it is treating draft-conditioned judgment as independent evidence. Blind Oracle separates derivation from extraction; the antipattern fuses them through a single judge prompt that anchors on the submission. This is the same body as **Comparator**'s fused QA evaluator antipattern from another angle. Comparator cares that extraction, comparison, and judgment are fused. Blind Oracle cares that the fused call conditions expected on the draft. ### Pattern: derive expected before reading the draft The structured implementation computes expected through a function that cannot read the draft, then extracts observed and compares. ```python from collections.abc import Callable from dataclasses import dataclass from typing import Literal Question = str Draft = str Expected = str Observed = str Verdict = Literal["pass", "fail"] DerivationSource = Literal["spec", "reference", "blind_rederivation"] @dataclass(frozen=True) class BlindVerification: question: Question expected: Expected observed: Observed verdict: Verdict derivation_source: DerivationSource extraction_method: str comparator: str draft_in_derivation_context: bool def blind_verify( question: Question, draft: Draft, derive_expected: Callable[[Question], tuple[Expected, DerivationSource]], extract_observed: Callable[[Draft], Observed], compare: Callable[[Expected, Observed], Verdict], ) -> BlindVerification: expected, source = derive_expected(question) observed = extract_observed(draft) verdict = compare(expected, observed) return BlindVerification( question=question, expected=expected, observed=observed, verdict=verdict, derivation_source=source, extraction_method="programmatic", comparator=getattr(compare, "__name__", type(compare).__name__), draft_in_derivation_context=False, ) def derive_expected_from_spec(question: Question) -> tuple[Expected, DerivationSource]: filing_spec = { "What is the named operator in this 10-K excerpt?": { "named_operator": "Eastern Pacific Holdings", }, } return filing_spec[question]["named_operator"], "spec" def extract_answer(draft: Draft) -> Observed: return draft.strip() def exact_match(expected: Expected, observed: Observed) -> Verdict: return "pass" if expected == observed else "fail" question = "What is the named operator in this 10-K excerpt?" correct_draft = "Eastern Pacific Holdings" adversarial_draft = "Western Atlantic Holdings" correct_result = blind_verify( question, correct_draft, derive_expected_from_spec, extract_answer, exact_match, ) adversarial_result = blind_verify( question, adversarial_draft, derive_expected_from_spec, extract_answer, exact_match, ) assert ( correct_result.expected == adversarial_result.expected and correct_result.verdict == "pass" and adversarial_result.verdict == "fail" ) ``` The load-bearing move is the call order and function boundary. `derive_expected_from_spec` has no draft parameter, so the draft cannot set expected. The draft is read only by `extract_answer`, which produces observed for comparison. Anthropic's outcome-grader notebook is the closest OSS instance. It runs the grader as a second agent with its own context window, cannot see the writer's reasoning, and re-reads the artifact against a rubric. It is adjacent rather than pure Blind Oracle because the grader does not derive a separately channelled expected value before grading. It still operationalizes the key separation: the grader's evidence path is structurally independent of the writer's evidence path. Outcome grading also composes with **Admissibility Gate**. Blind Oracle separates the derivation channel; Admissibility Gate defines what evidence is admissible for acceptance. ## Determinism Move Blind Oracle constrains `context_contamination` by separating the derivation channel from the production channel. Expected comes from spec, reference, or re-derivation. Observed comes from the draft. The two never share an input path. It also constrains `self_review_bias` as a design judgment for same-context systems: when the writer's draft is already in the verifier's context, the check can inherit the writer's anchors before it has independent evidence. The determinism move is enforced derivation independence; if the function that produces expected can read the draft, the verification is anchored regardless of what the prompt says. ## Observable Signal Every Blind Oracle report should include: * question or spec input; * expected value; * derivation source (`spec`, `reference`, `blind_rederivation`); * observed value; * extraction method (`programmatic`, `model_mediated`); * comparator name; * verdict; * draft-in-derivation-context boolean. A useful report makes the independence visible: ```text question: "What is the named operator in this 10-K excerpt?" expected: "Eastern Pacific Holdings" derivation_source: spec observed: "Eastern Pacific Holdings" extraction_method: programmatic comparator: exact_match verdict: pass draft_in_derivation_context: false ``` In the code sample, `draft_in_derivation_context: false` records the wrapper's call shape: `derive_expected` received the question and no draft argument. The paired correct and adversarial runs are the behavior check that expected stays stable while the verdict changes. ## Failure Modes * **Submission-First Prompting:** the verifier prompt places the draft in context before asking for a verdict. The model's expected becomes the draft restated. Derive expected first in a separate call that does not receive the draft. * **Fused Extraction-Comparison-Judgment:** one model call extracts observed, compares it to expected, and renders a verdict. Split the task into deterministic extract and named compare operators. * **Reasoning-Trace Leakage:** the verifier does not see the draft but receives the writer's chain-of-thought or scratchpad. Treat reasoning trace and shortcut history as draft-equivalent inputs. * **Reference-As-Adornment:** a reference value is supplied alongside the submission in one prompt. Hide the submission until expected is derived from the reference. ## Use When Use this pattern when: * the property under check has a derivable expected value, such as math, parses, structural counts, or executable checks; * false positives from draft-anchored judgment are a known failure mode; * the spec or reference can be processed independently of the draft; * generation and verification share a context window; * the verifier is LLM-based and could otherwise be conditioned on the draft. ## Do Not Use When Do not reach for Blind Oracle when: * expected cannot be derived without reading the draft; * **Executable Analog** can specialize the pattern with compilation, execution, or runtime traces; * no expected exists and the right pattern is **Admissibility Gate** or a calibrated **Judge Harness**; * the artifact and the spec are the same object and derivation must reference the artifact. If derivation independence cannot be enforced, label the verification as draft-anchored and escalate to a Judge Harness with perturbation, repetition, and calibration. ## Evidence * **Verification Design Principle 2:** the design doc names independence between generation and verification and warns that a verifier conditioned on the original output copies the same errors. * **Chain-of-Verification:** the research callout records the factored and revised variant, where verification questions are answered without conditioning on the draft, and reports the FACTSCORE improvement from 55.9 to 71.4 [arXiv:2309.11495]. * **[LangChain](https://github.com/langchain-ai/langchain) criteria and scoring evaluators:** the verification sweep records prompt shapes where criteria, submission, and sometimes reference are shown together before the model renders judgment. * **LangChain QA evaluator:** the same sweep records a fused extraction, comparison, and judgment path shared with the Comparator antipattern. * **Anthropic outcome grader:** the verification sweep records a supporting instance: a separate grader context inspects the artifact against a rubric without seeing the writer's reasoning. ## Related Patterns * **Comparator:** provides the named operator that decides verdict from `expected` and `observed`. * **Admissibility Gate:** defines admissibility for acceptance; Blind Oracle defines the upstream derivation channel. * **Cross-Family:** addresses which model verifies; Blind Oracle addresses what the verifier is allowed to see. * **Executable Analog:** is the executable specialization, and strongest form, of Blind Oracle when expected can be derived by executing the spec. * **Judge Harness:** wraps a Blind Oracle judge with perturbation, repetition, and calibration when the derivation channel still has subjective slack. --- Canonical: https://verificationdesign.com/patterns/verification/comparator/ Source: ai-design-patterns/cards/comparator.md # Comparator *(Verification Pattern)* ## Name **Comparator** Also known as: Named Comparison, Comparison Operator, Match Mode. ## Intent Express verification comparison as a named operator from a finite family, so the verdict is a deterministic function of `(expected, observed, operator, threshold, normalization)` rather than a model's interpretation of "does this look right?" ## Problem Many verification steps hide the actual comparison inside an LLM prompt: * "Here is the question, the student's answer, and the true answer. Grade CORRECT or INCORRECT." * "Here is the expected behavior and the actual behavior. Does the output satisfy the requirement?" * "Here is the rubric and the answer. Score it." Those prompts fuse three operations that should be separate: 1. extract the observed value, 2. compare it to the expected value, 3. produce the verdict. Once those operations are fused, the recorded result is only the model's label. A reviewer cannot replay the comparison without rerunning the judge. They cannot see what tolerance was applied, whether whitespace mattered, whether order mattered, whether JSON was canonicalized, or whether the model silently used a semantic standard that was never specified. `verification_design.md` Principle 5 says verifiers need explicit criteria: what is checked, what is expected, and how it is checked. Comparator names the "how." Principle 6 then pushes the same idea toward executable verification: once the comparison is named and deterministic, the verdict can be replayed. ## Forces * **Exact vs. semantic equivalence.** Character-for-character equality is auditable, but many tasks need normalization, regex matching, structured comparison, or trajectory matching. * **Strict vs. tolerant.** The system must decide whether casing, punctuation, ordering, extra fields, or near misses matter. * **Single value vs. structured state.** A comparator may operate on one string, a JSON tree, a list of tool calls, or a predicate set. * **Operator design vs. judge call.** A named operator costs more to design than a one-off prompt, but it is cheaper to run and easier to audit. * **Replayable verdict vs. opaque label.** A comparator report shows the operator and values; a fused judge report shows only a label. ## Solution Maintain a finite family of named comparison operators. For each verification step, choose the least restrictive operator allowed by the stated criterion, rather than tuning the operator until the current answer passes. Extract the observed value before comparison. Do not let the comparison operator fetch state, parse a model's reasoning, or decide what should be checked. It receives `expected`, `observed`, and operator settings. It returns a structured record with the operator name, comparison values, normalization, notes, score, threshold, and derived verdict. Common shapes: * **String comparators:** exact match, normalized exact match, regex match, JSON canonical equality, JSON edit distance. * **Structured-event comparators:** trajectory exact, trajectory in-order, trajectory any-order. * **Predicate aggregators:** per-criterion predicates with a named aggregation policy such as all-must-pass or k-of-n. If no named operator fits, that is useful information. Tighten the answer format so a comparator applies, or escalate to a **Judge Harness**. Do not silently collapse the check into an inline LLM judgment. ## Mechanism 1. **Define the comparison surface.** Name whether the comparison is over a string, regex, JSON tree, structured event sequence, or predicate set. 2. **Pick a named operator.** Choose `exact`, `normalized_exact`, `regex`, `json_canonical`, `json_distance`, `trajectory_exact`, `trajectory_in_order`, `trajectory_any_order`, or a named aggregator. 3. **Extract observed value separately.** Extraction happens before the comparator runs and is recorded verbatim. 4. **Apply the operator.** Return a structured score with `{operator, expected, observed, score, threshold, normalization}`; for `regex`, the expected field is the pattern being matched. 5. **Record every comparison.** Reports include operator, expected, observed, normalization, notes, score, threshold, and a verdict derived from `score >= threshold` for passes and failures. ## Pattern / Antipattern The same task: verify whether a student's answer matches a known answer. The antipattern asks an LLM to perform the comparison and emit a label. The pattern applies a named comparator to extracted values and records the full comparison. ### Antipattern: fused comparison prompt The naive implementation puts question, submitted answer, and true answer in one prompt. Extraction, comparison, and judgment happen inside the model call. ```python def grade_answer(question: str, student_answer: str, true_answer: str, model) -> dict: prompt = f""" You are a teacher grading a quiz. QUESTION: {question} STUDENT ANSWER: {student_answer} TRUE ANSWER: {true_answer} Grade the student answer as CORRECT or INCORRECT. """ label = model.complete(prompt).strip().upper() passed = label.startswith("CORRECT") return { "verdict": "pass" if passed else "fail", "label": label, } ``` This may be a practical fallback for semantic equivalence, but it is not a comparator. The report does not say what was compared, which operator was used, what tolerance applied, or what observed value the comparison consumed. The model's final label is the only artifact. LangChain's QA evaluator evidence has this shape: the evaluator prompt includes the question, student answer, and true answer, and an LLM chain parses CORRECT or INCORRECT into a score. That is the right tool only when a named comparison operator would over-reject. It becomes the Comparator antipattern when exact, regex, JSON, trajectory, or execution comparison would have sufficed. ### Pattern: named operator dispatch The structured implementation separates extraction from comparison and makes the operator explicit. ```python import json import re from dataclasses import dataclass, field from typing import Any, Callable @dataclass(frozen=True) class ComparisonResult: operator: str expected: Any observed: Any score: float threshold: float normalization: list[str] notes: list[str] = field(default_factory=list) @property def passed(self) -> bool: return self.score >= self.threshold @property def verdict(self) -> str: return "pass" if self.passed else "fail" def to_report(self) -> dict[str, Any]: return { "operator": self.operator, "expected": self.expected, "observed": self.observed, "normalization": self.normalization, "notes": self.notes, "score": self.score, "threshold": self.threshold, "verdict": self.verdict, } def normalize_text(value: str) -> str: return " ".join(value.lower().strip().split()) def exact(expected: str, observed: str) -> tuple[float, list[str], list[str]]: return (1.0 if expected == observed else 0.0, [], []) def normalized_exact(expected: str, observed: str) -> tuple[float, list[str], list[str]]: return ( 1.0 if normalize_text(expected) == normalize_text(observed) else 0.0, ["lowercase", "strip", "collapse_whitespace"], [], ) def regex(expected: str, observed: str) -> tuple[float, list[str], list[str]]: try: matched = re.fullmatch(expected, observed) except re.error: return (0.0, [], ["invalid_regex_pattern"]) return (1.0 if matched else 0.0, [], []) def json_canonical(expected: str, observed: str) -> tuple[float, list[str], list[str]]: try: expected_json = json.dumps(json.loads(expected), sort_keys=True) observed_json = json.dumps(json.loads(observed), sort_keys=True) except json.JSONDecodeError: return (0.0, [], ["json_parse_failed"]) return (1.0 if expected_json == observed_json else 0.0, ["json_sort_keys"], []) OPERATORS: dict[str, Callable[[str, str], tuple[float, list[str], list[str]]]] = { "exact": exact, "normalized_exact": normalized_exact, "regex": regex, "json_canonical": json_canonical, } def compare(operator: str, expected: str, observed: str, threshold: float = 1.0) -> ComparisonResult: if operator not in OPERATORS: raise ValueError(f"Unknown comparator: {operator}") if not 0.0 < threshold <= 1.0: raise ValueError("threshold must be in (0, 1]") score, normalization, notes = OPERATORS[operator](expected, observed) return ComparisonResult( operator=operator, expected=expected, observed=observed, score=score, threshold=threshold, normalization=normalization, notes=notes, ) report = compare( operator="normalized_exact", expected="The answer is 42.", observed="the answer is 42.", ) assert report.passed is True assert compare("exact", "The answer is 42.", "the answer is 42.").passed is False assert report.normalization == ["lowercase", "strip", "collapse_whitespace"] malformed_json = compare("json_canonical", '{"answer": 42}', '{"answer":') assert malformed_json.passed is False assert malformed_json.notes == ["json_parse_failed"] assert report.to_report() == { "operator": "normalized_exact", "expected": "The answer is 42.", "observed": "the answer is 42.", "normalization": ["lowercase", "strip", "collapse_whitespace"], "notes": [], "score": 1.0, "threshold": 1.0, "verdict": "pass", } ``` LangChain's non-LLM evaluator family demonstrates this pattern at framework level: exact match, regex match, and JSON edit distance take reference and prediction values and return mechanical scores. Its evaluator schema distinguishes named non-LLM evaluators from LLM evaluators, which is the architectural distinction Comparator needs. ADK's `TrajectoryEvaluator` shows the structured-event variant. It compares actual and expected tool-call sequences using named match modes: `EXACT`, `IN_ORDER`, and `ANY_ORDER`. This illustrates Comparator as a finite family of declared tolerance contracts, not just equality. ## Determinism Move Comparator constrains `judge_subjectivity` by making the verdict a deterministic function of expected value, observed value, operator, threshold, and normalization. It constrains `criteria_drift` because a named operator is stable across runs in a way that a prompt-based judge's interpretation is not. The same `trajectory_in_order` operator should return the same verdict today and next week. If the operator changes, the change is a code or configuration diff, not a hidden shift in a judge prompt's interpretation. ## Observable Signal Every Comparator report should include: * operator name; * expected value or reference; * observed value extracted before comparison; * normalization steps applied before comparison; * notes for recorded parse or pattern failures; * score and threshold; * pass/fail verdict. A useful report is replayable from the record plus the pinned comparator implementation: ```text operator: normalized_exact expected: "The answer is 42." observed: "the answer is 42." normalization: lowercase, strip, collapse_whitespace notes: [] score: 1.0 threshold: 1.0 verdict: pass ``` ## Failure Modes * **Fused Judgment:** Extraction, comparison, and verdict collapse into one LLM prompt. Extract observed value first, then apply a named operator outside the model call. * **Wrong Operator:** The check uses `exact` when `regex`, `json_canonical`, or `trajectory_any_order` matches the criterion. Apply the Solution rule: choose the least restrictive operator the criterion permits. * **Hidden Normalization:** The operator lowercases, strips punctuation, or canonicalizes JSON without reporting that step. Record normalization and normalized comparison surfaces. * **No Tolerance Policy:** Structured events are compared without declaring exact, in-order, any-order, distance, or threshold semantics. Reject `compare(expected, observed)` without a mode. ## Use When Use this pattern when: * the check has a known expected value, pattern, reference object, or expected event sequence; * the observed value can be extracted separately from the comparison; * a named operator covers the comparison or can be defined cheaply; * the report must be auditable without rerunning a model; * the same comparison will run repeatedly in CI, regression tests, or agent loops. ## Do Not Use When Do not reach for Comparator when: * the answer is genuinely open-ended and no reference exists; * semantic equivalence is the actual task and a tight comparator would reject acceptable answers; * designing a comparator costs more than a one-off human review; * the comparison needs a calibrated model judge. Use **Judge Harness** instead of an inline fused judgment. ## Evidence * **Verification Design Principles 5 and 6:** The design doc requires explicit criteria and prefers executable verification. Comparator names the comparison method inside that criteria record. * **[LangChain](https://github.com/langchain-ai/langchain) non-LLM evaluators:** The evidence summary records exact, regex, and JSON distance evaluators that take reference and prediction as inputs and return mechanical scores. * **LangChain evaluator schema:** The framework distinguishes named non-LLM evaluator types from LLM evaluators, making the operator family explicit. * **[ADK](https://github.com/google/adk-python) TrajectoryEvaluator:** The evidence summary records `EXACT`, `IN_ORDER`, and `ANY_ORDER` modes for structured tool-call trajectories. * **LangChain QA evaluator:** The verification sweep records the fused-comparison antipattern: question, answer, and prediction go into one LLM prompt, and CORRECT or INCORRECT is parsed as the score. ## Related Patterns * **Executable Analog:** extracts the observed value; Comparator names the operator that compares it. * **Blind Oracle:** fused comparison often violates blind evaluation because the judge sees the draft beside the truth. * **Judge Harness:** handles cases where no named comparator fits, but only with calibration and reliability checks. * **Constitution:** can name expected values and accepted comparator operators. * **Delta:** produces relative observed values that Comparator can evaluate with a declared numeric tolerance operator. --- Canonical: https://verificationdesign.com/patterns/verification/delta/ Source: ai-design-patterns/cards/delta.md # Delta *(Verification Pattern)* ## Name **Delta** Also known as: Baseline Assertion, Relative State Verification, Ambient Isolation. ## Intent Check whether the observed change in environment state satisfies an action's expected postcondition; attribution to that action requires isolation or a propagated causal tag. ## Problem Agentic systems often operate in shared, messy, or long-lived environments (databases, CI/CD runners, live cloud accounts). When a verification step checks for an absolute threshold (e.g., `total_flows >= 5` or `task_status == "completed"`), it is vulnerable to ambient state contamination. If the system already had 5 flows before the agent acted, the assertion passes trivially. The verification system creates a green report, but it proved absolutely nothing about the agent's actions. Because agents can rationalize a coincidental pass as "my action succeeded," they are blind to the fact that their tools actually failed or were never called. ## Forces * **Shared vs. Ephemeral Environments.** Spinning up a perfectly clean, mocked environment for every agent action is computationally expensive and sometimes impossible; operating in dirty environments requires careful state isolation. * **Absolute vs. Relative logic.** It is easier to write `assert count == 10` than to capture a pre-state, pass it through the workflow, and write `assert post_count - pre_count == 10`. ## Solution Record a state baseline immediately before the agent acts. After the agent acts, measure the state again. The verification criteria must assert only on the mathematically or logically computed *delta* between the post-state and the pre-state. If the agent is supposed to create a new user, assert `user_count_post - user_count_pre == 1`. If the agent is supposed to append a log line, assert `len(log_lines_post) - len(log_lines_pre) == 1`. ## Mechanism 1. **Pre-Hook (Baseline Capture):** Query the specific environment metric before handing control to the agent. 2. **Context Passing:** Store the baseline alongside the current execution trajectory and treat it as immutable once captured. 3. **Agent Execution:** The agent performs its task. 4. **Post-Hook (Measurement):** Query the exact same environment metric. 5. **Delta Calculation:** Subtract, diff, or otherwise compare the two states, and assert on the diff. ## Pattern / Antipattern The same task: verify that an agent created one new flow in a shared environment. ### Antipattern: absolute assertion The naive implementation checks the final count only. It can pass before the agent does anything. ```python def verify_flow_created_naive(get_flow_count): observed = get_flow_count() return { "passed": observed >= 1, "observed": observed, "expected": "at least 1 flow", } ``` ### Pattern: baseline plus delta assertion The structured implementation captures the baseline before the agent acts and verifies the change measured during this run's verification window. ```python from typing import Callable class DeltaVerifier: def __init__(self, metric_name: str, fetch_state_fn: Callable[[], int], expected_delta: int): self.metric_name = metric_name self.fetch_state_fn = fetch_state_fn self.expected_delta = expected_delta self.pre_state = None self.post_state = None def capture_baseline(self) -> int: """Called by the orchestrator before the agent acts.""" self.pre_state = self.fetch_state_fn() return self.pre_state def verify_delta(self) -> dict: """Called by the verifier after the agent acts.""" if self.pre_state is None: raise ValueError("baseline never captured; cannot verify a delta") self.post_state = self.fetch_state_fn() actual_delta = self.post_state - self.pre_state return { "check": f"{self.metric_name}_delta", "metric": self.metric_name, "passed": actual_delta == self.expected_delta, "expected_delta": self.expected_delta, "actual_delta": actual_delta, "pre_state": self.pre_state, "post_state": self.post_state, } # A shared environment that already holds unrelated flows before the agent runs. flows = [{"id": 1}, {"id": 2}, {"id": 3}] def get_flow_count() -> int: return len(flows) def verify_flow_created_naive_local(get_flow_count: Callable[[], int]) -> dict: observed = get_flow_count() return { "passed": observed >= 1, "observed": observed, "expected": "at least 1 flow", } flow_verifier = DeltaVerifier("active_flows", get_flow_count, expected_delta=1) flow_verifier.capture_baseline() # pre-action: records 3 assert verify_flow_created_naive_local(get_flow_count)["passed"] is True flows.append({"id": 4}) # the agent creates exactly one flow this run report = flow_verifier.verify_delta() # post-action check assert report == { "check": "active_flows_delta", "metric": "active_flows", "passed": True, "expected_delta": 1, "actual_delta": 1, "pre_state": 3, "post_state": 4, } assert report["pre_state"] == 3 and report["post_state"] == 4 assert report["passed"] # the measured change in this window is +1 ``` Aider's command tests show this shape in a real suite: they capture `initial_count = len(coder.abs_fnames)` before dropping files from a chat session, then after the drop assert that membership changed and `len(coder.abs_fnames) == initial_count - 1`. The check is on the relative file-set change, not an absolute count, so it holds no matter how many files were already in the session. ## Determinism Move Delta constrains `ambient_state` by turning an absolute assertion into a relative assertion scoped to the current run. The baseline and post-state expose whether the matching object was already present before the agent acted. The determinism move is scoping the assertion to a captured baseline, so a pass means the measured state changed by the expected amount during this run's verification window. Full causal attribution under concurrent activity requires scoped metrics, exclusive access, or pairing with **Causal Tag**. ## Observable Signal The observable signal reports the baseline, the new state, the computed delta, the expectation, and the metric being measured. Every report should include: * the baseline value captured before the agent acted; * the post-action value captured after the agent acted; * the computed delta; * the expected delta; * the named metric used for both captures. ```text check: active_flows_delta metric: active_flows passed: true pre_state: 3 post_state: 4 expected_delta: 1 actual_delta: 1 ``` ## Failure Modes * **Concurrency Contamination:** If other agents or users are modifying the environment at the exact same time, the delta might be +2 instead of +1. Mitigation: combine Delta with **Causal Tag**, tagging the agent's specific artifacts with a UUID. * **Stale Baselines:** Capturing the baseline too early (e.g., at system boot instead of immediately before the specific step being verified). ## Use When Use this pattern when: * the environment is shared, persistent, or carries non-trivial pre-existing state; * the metric being measured exists in the environment before the agent acts; * concurrent activity is possible and the metric can be scoped, locked, or paired with Causal Tag; * the assertion needs to prove the state changed during the agent's verification window, with causal attribution handled by scoping, exclusive access, or Causal Tag when concurrency is present. ## Do Not Use When Do not reach for Delta when: * the environment is fully ephemeral (fresh container per run, mocked DB); * the metric does not exist pre-action and only the absolute post-state is meaningful; * you cannot capture a baseline before the agent acts (e.g., the agent is invoked black-box); * the action is destructive or non-replayable and a pre-state read would alter the outcome. In those cases, absolute assertions or Causal Tag are simpler and more honest. ## Evidence * **[Aider](https://github.com/Aider-AI/aider) command tests:** the delta sweep records a direct instance in aider's `test_commands.py`: the drop-file tests capture `initial_count = len(coder.abs_fnames)` and then assert `len(coder.abs_fnames) == initial_count - 1`, a relative file-set delta rather than an absolute count. * **Test Flakiness Analysis:** Luo et al. identify state contamination and order dependence as recurring causes of flaky tests in software engineering (Luo et al., FSE 2014). * **Verification Design Principles (Principle 9):** "Assertions must prove the system's actions caused the expected outcome, not that the environment happened to already contain matching data." ## Related Patterns * **State Baseline:** the Context pattern that captures and carries the pre-state Delta asserts on. * **Causal Tag:** complementary for shared or concurrent environments; tag the agent's artifacts so a +2 from another actor does not mislead the delta. * **Comparator:** Delta is one named comparison operator (relative change) in the broader operator family. * **Executable Analog:** a delta assertion is the executable check for a state-change claim. * **Backpressure:** a failed delta check is a concrete failure signal that can drive a bounded rerun. --- Canonical: https://verificationdesign.com/patterns/verification/judge-harness/ Source: ai-design-patterns/cards/judge-harness.md # Judge Harness *(Verification Pattern)* ## Name **Judge Harness** Also known as: Judge Reliability Harness, LLM-as-Judge Harness, Calibrated Judge Wrapper, Perturbation + Repetition Layer. ## Intent Wrap an LLM judge in a structural harness of perturbation, repetition, calibration, and reporting so that one judge verdict becomes a measured signal with visible consistency and bias controls. The harness is the layer around the judge. It is not the judge prompt itself. A prompt that tells the judge to avoid bias may help the call, but it does not create the harness. ## Problem An LLM judge returns one verdict. The system treats that verdict as ground truth. That can look disciplined in code: * the judge prompt warns the model to avoid position, length, and name biases; * assistant A and assistant B are shown in fixed labeled slots, then a single pairwise verdict is parsed; * the judge is repeated only when the verdict disagrees with the architect's expectation; * a numeric judge score is logged, but the judge is never calibrated against human labels; * downstream consumers receive only `PASS` or `FAIL`, not the sample distribution behind it. The verifier failure is not that the judge is an LLM. The failure is treating a single subjective sample as a reliable measurement. Prompt-level caution does not expose whether the judge is stable under harmless input changes, whether repeated samples agree, or whether aggregate verdicts track human labels. `verification_design.md` Principle 7 names the broader boundary: cross-family beats self-verification. The Judge Reliability Harness update narrows it: cross-family LLM judges are not automatically reliable verification signals; they need perturbation tests, observed-behavior reporting, and calibration [arXiv:2603.05399]. The Gaming the Judge update extends the caution: even cross-family judges can be manipulated through chain-of-thought, so the harness must report observed judge behavior under perturbation, not only raw verdicts [arXiv:2601.14691]. ## Forces * **Single verdict vs. distribution of verdicts.** One call is cheap; a distribution exposes variance the single call hides. * **Prompt-level mitigation vs. structural mitigation.** "Avoid position bias" in the prompt is a wish; swapping positions before each run is a property. * **Cost vs. confidence.** Each perturbation, swap, or repetition multiplies the judge call budget; the harness trades budget for measured reliability. * **Calibration anchor vs. drift.** Without a human-label calibration set, harness output is internally consistent but externally unanchored. * **Aggregation rule vs. ambient majority.** Majority vote is one rule; abstain-on-disagreement is another. The harness names the rule. * **Reporting boundary.** A harness that aggregates internally but reports only a final verdict loses the calibration evidence the downstream consumer needs. ## Solution Put a structural harness around the judge call. The judge prompt can still say "be careful about position bias," but the harness enforces the check outside the prompt. The harness has four layers: * **Perturbation:** vary judge inputs in ways that should not change the verdict, such as paraphrase, formatting changes, name swap, or position swap when the judge compares two answers. * **Repetition:** invoke the judge N times per perturbation with independent samples. * **Aggregation rule:** choose a named rule that turns the sample into a harness verdict, such as majority, supermajority, or abstain-on-disagreement. * **Calibration and reporting:** compare aggregated verdicts against a human-labeled calibration set, then report the verdict, sample distribution, consistency rate, and calibration anchors together. A bias-warning prompt is a judge input. Perturbation, repetition, aggregation, calibration, and reporting are harness properties. ## Mechanism 1. **Identify the judge call surface.** Find the LLM call that returns the verdict. 2. **Define the perturbation set.** Name which input variations should preserve the verdict. 3. **Set the repetition count and aggregation rule.** Choose N and the rule before seeing the result. 4. **Maintain calibration labels.** Keep a human-labeled calibration set and periodically re-calibrate the judge. 5. **Stamp the verdict.** Record sample distribution, consistency rate, perturbations applied, aggregation rule, and calibration source. ## Pattern / Antipattern The two examples use different judge shapes to show the same reliability boundary. The antipattern is a one-call pairwise judge that trusts one A/B/C parse. The pattern is a single-answer PASS/FAIL judge wrapped as an instrument with sampling, perturbation, aggregation, calibration, and reportable behavior. ### Antipattern: one-call judge with bias warning The naive implementation places the bias instruction in the prompt and treats the first parsed verdict as the result. ```python from typing import Literal Verdict = Literal["A", "B", "C"] def one_call_pairwise_judge(model, question: str, answer_a: str, answer_b: str) -> Verdict: prompt = f""" You are an impartial judge. Avoid position bias, length bias, and name bias. [Question] {question} [Assistant A] {answer_a} [Assistant B] {answer_b} Return exactly one verdict: [[A]], [[B]], or [[C]]. """ raw = model.complete(prompt) if "[[A]]" in raw: return "A" if "[[B]]" in raw: return "B" return "C" ``` The prompt warns about bias, but the code never swaps positions, repeats the call, calibrates against labels, or reports the sample distribution. One parse becomes the verdict. LangChain's pairwise evaluator is the canonical antipattern pick: it puts explicit position, length, and name bias warnings in the prompt, fixes assistant A and assistant B in labeled slots, then parses one pairwise verdict, `[[A]]`, `[[B]]`, or `[[C]]`, from a single judge call. The LangChain scoring evaluator is the sibling shape: an impartial one-shot prompt returns a 1-to-10 rating instead of an A/B/C choice. A bias-warning in the judge prompt is not wrong on its own; it is a useful nudge. It becomes the Judge Harness antipattern when the bias-warning is treated as the harness, with no perturbation, repetition, calibration, or reporting around the call. This also cross-lists with **Blind Oracle** when the submitted answer is already in the verifier's context before independent expected evidence exists. ### Pattern: calibrated perturbation and repetition wrapper The structured implementation wraps the judge call and returns the measurement boundary with the verdict. ```python from collections import Counter from collections.abc import Callable, Iterable from dataclasses import dataclass from typing import Literal Verdict = Literal["PASS", "FAIL"] AggregationRule = Literal[ "majority", "supermajority", "abstain_on_disagreement", ] SUPERMAJORITY_THRESHOLD = 2 / 3 @dataclass(frozen=True) class PromptInputs: question: str answer: str @dataclass(frozen=True) class Perturbation: name: str apply: Callable[[PromptInputs], PromptInputs] @dataclass(frozen=True) class CalibrationSet: name: str labeled_examples: tuple[tuple[PromptInputs, Verdict], ...] @dataclass(frozen=True) class HarnessResult: verdict: Verdict | Literal["ABSTAIN"] sample_distribution: dict[str, int] consistency_rate: float perturbations: tuple[str, ...] repetitions: int aggregation_rule: AggregationRule calibration_source: str judge_model: str calibrated_precision: float | None calibrated_recall: float | None class JudgeHarness: def __init__( self, perturbations: Iterable[Perturbation], repetitions: int, aggregation_rule: AggregationRule, calibration_set: CalibrationSet, ): self.perturbations = tuple(perturbations) if not self.perturbations: raise ValueError("JudgeHarness requires at least one perturbation") if repetitions <= 0: raise ValueError("JudgeHarness repetitions must be positive") self.repetitions = repetitions self.aggregation_rule = aggregation_rule self.calibration_set = calibration_set def evaluate( self, judge_call: Callable[[PromptInputs], Verdict], prompt_inputs: PromptInputs, judge_model: str, ) -> HarnessResult: samples: list[Verdict] = [] for perturbation in self.perturbations: perturbed = perturbation.apply(prompt_inputs) for _ in range(self.repetitions): samples.append(judge_call(perturbed)) counts = Counter(samples) verdict, count = counts.most_common(1)[0] total = sum(counts.values()) consistency_rate = count / total if self.aggregation_rule == "abstain_on_disagreement" and len(counts) > 1: final_verdict: Verdict | Literal["ABSTAIN"] = "ABSTAIN" elif self.aggregation_rule == "supermajority" and count / total < SUPERMAJORITY_THRESHOLD: final_verdict = "ABSTAIN" else: final_verdict = verdict precision, recall = calibrate(judge_call, self.calibration_set) return HarnessResult( verdict=final_verdict, sample_distribution=dict(counts), consistency_rate=consistency_rate, perturbations=tuple(item.name for item in self.perturbations), repetitions=self.repetitions, aggregation_rule=self.aggregation_rule, calibration_source=self.calibration_set.name, judge_model=judge_model, calibrated_precision=precision, calibrated_recall=recall, ) def calibrate( judge_call: Callable[[PromptInputs], Verdict], calibration_set: CalibrationSet, ) -> tuple[float, float]: labels = calibration_set.labeled_examples if not labels: return 0.0, 0.0 predicted_positive = 0 true_positive = 0 actual_positive = 0 for inputs, human_label in labels: prediction = judge_call(inputs) predicted_positive += int(prediction == "PASS") true_positive += int(prediction == "PASS" and human_label == "PASS") actual_positive += int(human_label == "PASS") precision = true_positive / predicted_positive if predicted_positive else 0.0 recall = true_positive / actual_positive if actual_positive else 0.0 return precision, recall def identity(inputs: PromptInputs) -> PromptInputs: return inputs class SequencedJudge: def __init__(self, verdicts: tuple[Verdict, ...]): self.verdicts = verdicts self.index = 0 def __call__(self, inputs: PromptInputs) -> Verdict: verdict = self.verdicts[self.index % len(self.verdicts)] self.index += 1 return verdict prompt_inputs = PromptInputs( question="Does the answer cite the required source?", answer="Yes. It cites the required source directly.", ) judge_sequence: tuple[Verdict, ...] = ( "PASS", "PASS", "PASS", "FAIL", "PASS", "FAIL", "PASS", "FAIL", "PASS", ) majority_harness = JudgeHarness( perturbations=[ Perturbation("paraphrase", identity), Perturbation("format_change", identity), ], repetitions=4, aggregation_rule="majority", calibration_set=CalibrationSet("human_labeled_set_v1", ((prompt_inputs, "PASS"),)), ) majority_result = majority_harness.evaluate(SequencedJudge(judge_sequence), prompt_inputs, "gpt-4o") assert majority_result.verdict == "PASS" assert majority_result.sample_distribution == {"PASS": 5, "FAIL": 3} supermajority_harness = JudgeHarness( perturbations=[ Perturbation("paraphrase", identity), Perturbation("format_change", identity), ], repetitions=4, aggregation_rule="supermajority", calibration_set=CalibrationSet("human_labeled_set_v1", ((prompt_inputs, "PASS"),)), ) supermajority_result = supermajority_harness.evaluate( SequencedJudge(judge_sequence), prompt_inputs, "gpt-4o", ) assert supermajority_result.verdict == "ABSTAIN" result = majority_harness.evaluate(SequencedJudge(judge_sequence), prompt_inputs, "gpt-4o") expected = HarnessResult( verdict="PASS", sample_distribution={"PASS": 5, "FAIL": 3}, consistency_rate=0.625, perturbations=("paraphrase", "format_change"), repetitions=4, aggregation_rule="majority", calibration_source="human_labeled_set_v1", judge_model="gpt-4o", calibrated_precision=1.0, calibrated_recall=1.0, ) assert result == expected ``` The load-bearing move is the returned measurement object. The downstream consumer can see that the verdict came from two perturbations, four repetitions per perturbation, majority aggregation, and a named human-labeled calibration source. ADK `llm_as_judge` plus `rubric_based_evaluator` is the closest OSS contrast. It already gives repetition and named majority-vote aggregation. It still needs the remaining harness layers: input perturbation, external calibration anchors, and explicit sample-distribution and consistency reporting. The returned `PerInvocationResult` exposes selected aggregate rubric scores and status, not a count or distribution. ## Determinism Move Judge Harness constrains `sampling_variance` by replacing one judge sample with an aggregated distribution. The verdict is no longer whichever sample happened to arrive first. It constrains `judge_subjectivity` by anchoring aggregate verdicts to a human-labeled calibration set. The judge's reading is checked against an external reference rather than accepted as its own proof. The determinism move is structural N-and-anchor; a single judge call without perturbation, repetition, or calibration is a verdict, not a measurement. ## Observable Signal Every Judge Harness report should include: * judge model identity; * perturbations applied, such as `paraphrase`, `format_change`, `name_swap`, `position_swap` for pairwise judges, or `none`; * repetitions per perturbation; * aggregation rule, such as `majority`, `supermajority`, or `abstain_on_disagreement`; * verdict; * sample distribution; * consistency rate; * calibration source; * calibrated precision and recall when calibration source is not `none`. A useful report exposes the measurement behind the verdict: ```text judge_model: gpt-4o perturbations: paraphrase, format_change repetitions_per_perturbation: 4 aggregation_rule: majority sample_distribution: 5 PASS, 3 FAIL verdict: PASS consistency_rate: 0.625 calibration_source: human_labeled_set_v1 calibrated_precision: 1.0 calibrated_recall: 1.0 ``` ## Failure Modes * **Bias-Warning As Harness:** the judge prompt says "avoid position bias" but the surrounding code does not swap positions. Replace the prompt-level wish with a structural pre-call perturbation. * **Single-Sample Verdict:** the harness runs one judge call and reports its verdict directly. Set `repetitions >= 2` and define an aggregation rule. * **Uncalibrated Aggregation:** the harness aggregates N samples but never compares the aggregated verdict against human labels. Maintain a calibration set; report precision and recall at the boundary. * **Hidden Distribution:** the harness aggregates internally but reports only the final verdict. Surface sample distribution and consistency rate in the verdict object. ## Use When Use this pattern when: * the verification path uses an LLM-based judge; * the verdict is high leverage, such as a gate for promotion, training, deployment, or automation; * false positives or false negatives are costly enough to justify multi-sample budget; * human labels exist or can be collected for a calibration set; * the judge family has documented bias modes that can be perturbed against. ## Do Not Use When Do not reach for Judge Harness when: * verification is fully executable, so an **Executable Analog** or **Comparator** can decide it without an LLM; * the verdict is low leverage and a single-sample judge call is proportionate; * no calibration set can be obtained and the harness would aggregate without an anchor; * the budget for repetitions is unavailable and the alternative is no judge at all. If a harness is infeasible, label the judge verdict as informal and report the single sample without aggregation framing. ## Evidence * **Verification Design Principle 7:** the design doc names cross-family verification as stronger than self-verification, while the Judge Reliability Harness update adds that LLM judges need perturbation tests, observed-behavior reporting, and calibration [arXiv:2603.05399]. * **Gaming the Judge update:** the design doc records that cross-family judges can still be manipulated through chain-of-thought, which makes observed judge behavior under perturbation part of the report [arXiv:2601.14691]. * **[LangChain](https://github.com/langchain-ai/langchain) pairwise evaluator:** the verification sweep records the canonical antipattern: explicit position, length, and name bias warnings in the prompt, fixed assistant A and assistant B slots, one parsed `[[A]]`, `[[B]]`, or `[[C]]` verdict, and no surrounding perturbation, repetition, calibration, or reporting code. * **LangChain scoring evaluator:** the same sweep records a sibling one-shot shape: an impartial scoring prompt returns a 1-to-10 rating with no surrounding harness machinery. * **[ADK](https://github.com/google/adk-python) LLM judge and rubric evaluator:** the verification sweep records partial harness machinery: multiple samples and majority-vote aggregation, but no inspected input perturbation, external calibration anchors, or explicit sample-distribution and consistency reporting. * **Anthropic cookbook knowledge-graph eval:** the OSS sweep records a mechanical metric paired with an explicit note about what it does not measure. That reporting-boundary discipline is part of the Judge Harness contract. ## Related Patterns * **Blind Oracle:** derives expected without conditioning on the draft; Judge Harness wraps the judge's verdict with structural N and calibration. * **Cross-Family:** addresses which model judges; Judge Harness addresses how the judge's output is sampled, aggregated, and anchored. * **Admissibility Gate:** defines admissibility for acceptance; Judge Harness measures the judge's reliability under perturbation. * **Constitution:** provides the shared rubric the harness can calibrate against. * **Comparator:** replaces Judge Harness when a deterministic comparison operator can decide the verdict without invoking an LLM judge. --- Canonical: https://verificationdesign.com/patterns/verification/admissibility-gate/ Source: ai-design-patterns/cards/admissibility-gate.md # Admissibility Gate *(Verification Pattern)* ## Name **Admissibility Gate** Also known as: Adversarial Frame (former name), Operationalized Skepticism, Evidence-First Verification, Default-No Rubric. ## Intent Replace tone-level skepticism instructions with admissibility rules that define what counts as proof, name common shortcut paths to reject, and invert the verifier's default from "accept if plausible" to "fail unless backed by trusted evidence." ## Problem LLM verifiers tend to approve plausible work, especially when they are asked confirmatory questions: * "Does this look right?" * "Review carefully." * "Be objective." * "Give constructive feedback, then approve when addressed." Those instructions describe a desired attitude, not a verification procedure. They ask the model to behave skeptically without defining what skepticism means. The judge may still treat the final answer as evidence, accept the agent's reasoning as if it were ground truth, or pass a lookalike source because it sounds close enough. `verification_design.md` Principle 4 names the frame shift: ask "what could fail?" rather than "does this look right?" The same principle cites SycEval's 58.19% sycophancy rate as a measured majority-rate finding, so agreement pressure is not a small edge case (AAAI 2025). A verifier that starts from plausibility will often rationalize acceptance. Admissibility Gate makes skepticism structural. It names what evidence is admissible, names what evidence is forbidden, lists common shortcuts to reject, and makes the default verdict `no` when evidence is missing. ## Forces * **Tone instruction vs. admissibility rule.** "Be critical" is a wish; "final answers and reasoning traces are not trusted evidence" is a constraint. * **Accept-if-plausible vs. reject-unless-supported.** Default verdicts matter. Missing evidence should fail the property, not defer the decision. * **Adversarial role vs. adversarial stance.** A critic role with a confirmatory prompt is not adversarial. * **Improvement loop vs. disconfirmation loop.** Evaluator-optimizer loops can refine outputs, but refinement is not the same as adversarial verification. * **Rubric cost vs. false approval cost.** Admissibility rules and shortcut lists take effort to write once; approving judges create recurring verification debt. ## Solution Write rubrics that operationalize skepticism. Do not rely on the verifier's tone. A rubric in this shape defines a compact contract: * what counts as trusted evidence; * what does not count as trusted evidence; * which domain shortcuts must be rejected; * what default verdict applies when evidence is missing; * the required output order: evidence first, rationale second, verdict last. Common shapes include evidence-admissibility rubrics, shortcut-rejection lists, and failure-hypothesis prompts. An evidence-admissibility rubric says trusted evidence must come from procedurally sound tool calls or verified external sources. Final answers, reasoning, summaries, interpretations, and flawed tool calls do not count. A shortcut-rejection list then names lookalikes to reject explicitly, such as an 8-K press-release exhibit when the rubric requires a 10-K or 10-Q filing. A failure-hypothesis prompt makes the verifier list ways the answer could be wrong before it can approve. The point is not to make the judge sound harsh. The point is to remove the judge's freedom to accept unsupported work. ## Mechanism 1. **Define admissibility.** Name the evidence categories that can support each property. 2. **Define forbidden evidence.** Name sources that cannot support the property, such as final answer text, reasoning trace, summaries, or flawed tool calls. 3. **Invert the default.** A property defaults to `no` unless admissible evidence supports it. 4. **Enumerate shortcuts.** List domain-specific substitutions the verifier must reject. 5. **Require evidence before verdict.** The report records evidence, then rationale, then yes/no verdict for each property. ## Pattern / Antipattern The same task: decide whether an agent's final answer satisfies a property. The antipattern labels a role "critic" but gives it a confirmatory prompt. The pattern makes admissible evidence load-bearing and defaults to rejection when evidence is missing. ### Antipattern: critic role with confirmatory prompt The naive implementation creates a critic agent but asks for general feedback and an approval token. There is no admissibility rule, no required failure search, no shortcut list, and no evidence-first report. ```python critic = AssistantAgent( name="critic", system_message=( "You are a critic. Provide constructive feedback. " "Respond with APPROVE if your feedback has been addressed." ), ) team = RoundRobinTeam( participants=[writer, critic], termination_token="APPROVE", ) ``` This is adversarial in role label only. It can improve an answer, but it does not require the critic to find disconfirming evidence before approval. The AutoGen Chainlit sample has this shape: a `critic` agent gives constructive feedback and the team terminates on `APPROVE`. This is sample code, not the library core, but examples are how patterns spread. A critic role with a confirmatory prompt can look like verification in code review while producing acceptance at runtime. Evaluator-optimizer loops need a different caveat. The Anthropic evaluator-optimizer notebook is a valid iteration pattern: the evaluator can return PASS, NEEDS_IMPROVEMENT, or FAIL and provide feedback for regeneration. The antipattern appears only when that improvement loop is misclassified as Admissibility Gate. "Has suggestions" is a weaker gate than "survived adversarial checks." Iteration and disconfirmation are separate patterns. ### Pattern: evidence admissibility with default-no verdict The structured implementation defines what evidence may support each property and refuses to approve when support is absent. ```python from dataclasses import dataclass from typing import Literal EvidenceType = Literal[ "procedurally_sound_tool_call", "verified_external_source", "final_answer", "reasoning_trace", "summary", "interpretation", "flawed_tool_call", ] Verdict = Literal["yes", "no"] ExclusionReason = Literal["forbidden", "inadmissible", "shortcut-rejected"] @dataclass(frozen=True) class Evidence: kind: EvidenceType detail: str source_label: str @dataclass(frozen=True) class PropertySpec: name: str admissible_evidence_types: tuple[EvidenceType, ...] forbidden_evidence_types: tuple[EvidenceType, ...] shortcut_rejections: tuple[str, ...] @dataclass(frozen=True) class ExcludedEvidence: detail: str reason: ExclusionReason def _join(values: tuple[str, ...]) -> str: return ", ".join(values) def _classify(item: Evidence, spec: PropertySpec) -> ExclusionReason | None: if item.kind in spec.forbidden_evidence_types: return "forbidden" if item.kind not in spec.admissible_evidence_types: return "inadmissible" if item.source_label in spec.shortcut_rejections: return "shortcut-rejected" return None def verify_property(evidence: list[Evidence], spec: PropertySpec) -> dict: admitted = [] excluded = [] for item in evidence: reason = _classify(item, spec) if reason is None: admitted.append(item.detail) else: excluded.append(ExcludedEvidence(detail=item.detail, reason=reason)) verdict: Verdict = "yes" if admitted else "no" rationale = ( "At least one admissible source supports the property." if admitted else "No admissible evidence supports this property." ) return { "property": spec.name, "admissible": list(spec.admissible_evidence_types), "forbidden": list(spec.forbidden_evidence_types), "shortcut_rejections": list(spec.shortcut_rejections), "evidence": admitted, "excluded": [ {"detail": item.detail, "reason": item.reason} for item in excluded ], "rationale": rationale, "default_verdict": "no", "verdict": verdict, } named_operator = PropertySpec( name="Named operator evidence comes from a 10-K or 10-Q filing", admissible_evidence_types=("verified_external_source",), forbidden_evidence_types=( "final_answer", "reasoning_trace", "summary", "interpretation", "flawed_tool_call", ), shortcut_rejections=("8-K press-release exhibit", "earnings recap", "news article"), ) final_answer_claim = Evidence( kind="final_answer", detail="Final answer says the 10-K confirms the operator.", source_label="final_answer", ) shortcut_source = Evidence( kind="verified_external_source", detail="8-K exhibit says the operator is named.", source_label="8-K press-release exhibit", ) ten_k_source = Evidence( kind="verified_external_source", detail="10-K filing says the operator is named.", source_label="10-K filing", ) strict_forbidden = verify_property([final_answer_claim], named_operator) assert strict_forbidden["verdict"] == "no" assert strict_forbidden["excluded"] == [ { "detail": "Final answer says the 10-K confirms the operator.", "reason": "forbidden", } ] strict_shortcut = verify_property([shortcut_source], named_operator) assert strict_shortcut["verdict"] == "no" assert strict_shortcut["excluded"] == [ { "detail": "8-K exhibit says the operator is named.", "reason": "shortcut-rejected", } ] strict_accept = verify_property([ten_k_source], named_operator) assert strict_accept["verdict"] == "yes" assert strict_accept["evidence"] == ["10-K filing says the operator is named."] lenient_kind_spec = PropertySpec( name=named_operator.name, admissible_evidence_types=("verified_external_source", "final_answer"), forbidden_evidence_types=("reasoning_trace", "summary", "interpretation", "flawed_tool_call"), shortcut_rejections=named_operator.shortcut_rejections, ) lenient_kind = verify_property([final_answer_claim], lenient_kind_spec) assert lenient_kind["verdict"] == "yes" assert lenient_kind["evidence"] == ["Final answer says the 10-K confirms the operator."] lenient_shortcut_spec = PropertySpec( name=named_operator.name, admissible_evidence_types=named_operator.admissible_evidence_types, forbidden_evidence_types=named_operator.forbidden_evidence_types, shortcut_rejections=(), ) lenient_shortcut = verify_property([shortcut_source], lenient_shortcut_spec) assert lenient_shortcut["verdict"] == "yes" assert lenient_shortcut["evidence"] == ["8-K exhibit says the operator is named."] def aggregate_reports(reports: list[dict]) -> dict: missing = sum(1 for report in reports if report["verdict"] == "no") return { "aggregate": "pass" if missing == 0 else "fail", "missing_evidence_count": missing, } def to_report(report: dict, aggregate: dict) -> str: excluded = "; ".join( f"{item['detail']} ({item['reason']})" for item in report["excluded"] ) return "\n".join( [ f"property: {report['property']}", f"admissible: {_join(tuple(report['admissible']))}", f"forbidden: {_join(tuple(report['forbidden']))}", f"shortcut_rejections: {_join(tuple(report['shortcut_rejections']))}", f"evidence: {report['evidence']}", f"excluded: {excluded}", f"rationale: {report['rationale']}", f"default_verdict: {report['default_verdict']}", f"verdict: {report['verdict']}", f"aggregate: {aggregate['aggregate']}", f"missing_evidence_count: {aggregate['missing_evidence_count']}", ] ) expected_report = """property: Named operator evidence comes from a 10-K or 10-Q filing admissible: verified_external_source forbidden: final_answer, reasoning_trace, summary, interpretation, flawed_tool_call shortcut_rejections: 8-K press-release exhibit, earnings recap, news article evidence: [] excluded: 8-K exhibit says the operator is named. (shortcut-rejected) rationale: No admissible evidence supports this property. default_verdict: no verdict: no aggregate: fail missing_evidence_count: 1""" assert to_report(strict_shortcut, aggregate_reports([strict_shortcut])) == expected_report ``` ADK's rubric-based final-response evaluator is the canonical evidence for this shape. It defines yes/no semantics, requires trusted evidence from procedurally sound tool calls, forbids deriving trusted evidence from the final answer, reasoning, summaries, interpretations, or flawed tool calls, and requires each property to output evidence, rationale, and verdict. The Anthropic outcome-grader notebook is the shortcut-list variant. Its rubric forces the grader to require concrete evidence and reject lookalike substitutions, including an 8-K press-release exhibit when the requirement was a 10-K or 10-Q. The Pattern code above makes that rejection executable by treating the source label as part of the rubric, so the line comes from the spec rather than the judge's taste. ## Determinism Move When the producer or a same-context verifier grades its own work, Admissibility Gate constrains `self_review_bias` by inverting the default from accept-if-plausible to reject-unless-supported. Missing evidence becomes a `no`, not an invitation to rationalize. It constrains `judge_subjectivity` by replacing tone instructions with admissibility rules and shortcut-rejection lists. The verifier no longer decides what counts as proof at runtime; the rubric says what counts. ## Observable Signal Every Admissibility Gate report should include: * admissible-evidence types for each property; * forbidden-evidence types for each property; * shortcut-rejection list; * evidence collected per property; * excluded evidence and its exclusion reason per property; * rationale per property; * default verdict when evidence is missing; * aggregate result with missing-evidence count. A useful report shows the rejection surface: ```text property: Named operator evidence comes from a 10-K or 10-Q filing admissible: verified_external_source forbidden: final_answer, reasoning_trace, summary, interpretation, flawed_tool_call shortcut_rejections: 8-K press-release exhibit, earnings recap, news article evidence: [] excluded: 8-K exhibit says the operator is named. (shortcut-rejected) rationale: No admissible evidence supports this property. default_verdict: no verdict: no aggregate: fail missing_evidence_count: 1 ``` ## Failure Modes * **Tone Instruction:** The rubric says "be objective" or "be skeptical" without defining admissible evidence. Replace stance instructions with evidence rules. * **Confirmatory Critic:** A role is named critic or adversary, but the prompt asks for constructive feedback and approval. Require evidence collection and a "why reject" field before any approval token can fire. * **Improvement Loop Misclassified:** Evaluator-optimizer iteration is treated as adversarial verification. Keep improvement and disconfirmation as distinct patterns. * **Hidden Shortcut:** The rubric lacks explicit rejection rules for common lookalikes. Enumerate shortcuts and make missing support fail by default. ## Use When Use this pattern when: * the verifier checks work produced by an LLM; * false approvals are more costly than false rejections; * the task has known shortcut paths or lookalike evidence; * generation and verification share a model family or context; * the rubric will be reused across many runs. ## Do Not Use When Do not reach for Admissibility Gate when: * the task is genuinely iterative and evaluator-optimizer is the desired loop; * the output is exploratory and no one has defined what "wrong" means yet; * strict admissibility would reject too many acceptable answers; * a named **Comparator** or **Executable Analog** can decide the check directly; * the right escalation is a calibrated **Judge Harness** around a subjective property. ## Evidence * **Verification Design Principle 4:** The design doc names adversarial framing and cites SycEval's sycophancy finding as the reason confirmatory review is unsafe (AAAI 2025). * **[ADK](https://github.com/google/adk-python) rubric-based final response quality:** The evidence summary records a direct pattern instance: trusted evidence must come from sound tool calls, forbidden evidence is listed, and missing support produces `no`. * **Anthropic outcome grader:** The verification sweep records a shortcut-rejection rubric that rejects an 8-K press-release exhibit when the requirement was a 10-K or 10-Q. * **[AutoGen](https://github.com/microsoft/autogen) Chainlit critic:** The antipattern cleanup sweep records a critic role whose prompt asks for constructive feedback and approval without required failure search. * **Anthropic evaluator-optimizer:** The same sweep records a valid improvement loop that becomes a partial-fit antipattern only when it is mistaken for adversarial verification. ## Related Patterns * **Adversary:** the orchestration role that applies this gate's admissibility and default-no logic through a mandatory negative channel. * **Blind Oracle:** derives expected evidence without conditioning on the draft; Admissibility Gate defines admissibility for acceptance. * **Comparator:** should replace Admissibility Gate when a named comparison operator can decide the check. * **Constitution:** can require admissibility rules and default-no posture as part of the criteria contract. * **Cross-Family:** reduces shared blind spots when the adversarial rubric is judged by a different model family. * **Judge Harness:** wraps the rubric with perturbation, repetition, calibration, and consistency checks. --- Canonical: https://verificationdesign.com/patterns/orchestration/cross-family/ Source: ai-design-patterns/cards/cross-family.md # Cross-Family *(Orchestration Pattern)* ## Name **Cross-Family** Also known as: Cross-Provider Verification, Independent-Family Judge, Family-Diverse Evaluation. ## Intent Run high-leverage generation and high-leverage assessment on deliberately different model families, and record both identities, so exposure to shared training-data biases and shared latent priors is reduced rather than assumed away, and any residual correlated error is attributable to named models. Cross-Family has two complementary roles. At verification time, the judge comes from a different family than the generator. At optimization time, the teacher or proposer comes from a different family than the model being shaped. Both roles serve the same mechanic: break same-family blind spots where one model's confidence becomes another model's evidence. ## Problem An LLM produces an answer. A second LLM is asked whether the answer looks right. Both are sampled from the same provider, often the same family, sometimes the same model. That can look independent in code: * one object is called `apprentice` and another object is called `grader`; * one prompt says "write" and another prompt says "review"; * one agent is a manager and another agent is a researcher; * the framework exposes a configurable judge slot. But the verification boundary may still collapse. The same blind spot that produced the error can also fail to detect it. Separate roles do not create independent evidence when the generator and verifier share the same model family. Common shapes: * one configured client is reused for both apprentice and grader; * separate clients are created, but both are `gpt-*`, both are `claude-*`, or both are the same Llama family; * role-diverse agents are wired together, but the judge defaults to whichever provider the framework imported first; * the verifier argument is left at `model=None`, and the framework silently chooses the default family. `verification_design.md` Principle 7 names the design rule: cross-family verification beats self-verification. The Judge Harness update narrows the claim: family diversity is necessary, not sufficient. A different-family judge still needs perturbation, repetition, and calibration around it. ## Forces * **Verification-time independence vs. operational simplicity.** One client per app is easy to configure; family-diverse verification adds providers, secrets, and routing. * **Provider availability vs. provider lock-in.** Verifier diversity assumes more than one provider is reachable. * **Cost vs. independence.** Cross-family verification pays for two providers; same-family verification pays for one. * **Role diversity vs. verifier diversity.** Different agent roles inside one family do not equal cross-family verification. The seam that matters is generator to verifier. * **Optimization-time vs. verification-time.** A teacher/student split can break same-family bias during prompt optimization; a different-family judge can break it during evaluation. * **Recorded identity vs. implicit identity.** A run that does not log generator and verifier identities cannot be audited for same-family bias. ## Solution At every leverage point where one model's output becomes another model's input as evidence, require that the two models come from different families, and stamp both identities into the artifact the verifier produces. "Different family" is a configured choice, not an accident of which client was imported first. The pattern lives at three layers: * **Configuration:** generator and verifier clients are constructed as separate objects, with explicit family or provider attributes. * **Routing:** the verification call site reads the verifier from configuration, not from a module-level default. * **Reporting:** the verdict object carries `generator_model`, `generator_family`, `verifier_model`, and `verifier_family`. A verdict that does not name both identities is not auditable. ## Mechanism 1. **Name the leverage points.** Identify calls where model output becomes load-bearing evidence: judge calls, optimization advice, debate adjudication, escalation arbitration. 2. **Construct distinct clients.** Generator and verifier clients are separate objects. Each carries provider and family metadata. 3. **Assert at routing time.** Before invoking the verifier, assert that `verifier.family != generator.family`. 4. **Carry both identities into the verdict.** The report contains generator identity and verifier identity. 5. **Wrap the verifier in a Judge Harness.** Cross-family is necessary, not sufficient. Consistency, perturbation, and calibration checks belong around the judge. ## Pattern / Antipattern The same task: ask a model to answer, then ask a model to grade the answer. The antipattern uses one family and treats the verdict as independent. The pattern requires family diversity at that boundary and records both identities. ### Antipattern: same-client grader The naive implementation stores one chat client on the grader and uses it for both answer extraction and correctness judgment. ```python from dataclasses import dataclass @dataclass(frozen=True) class ChatClient: model: str family: str def complete(self, prompt: str) -> str: ... class Grader: def __init__(self, client: ChatClient): self.client = client def extract_answer(self, transcript: str) -> str: return self.client.complete( "Extract the apprentice's final answer:\n" + transcript ) def judge_correctness(self, question: str, answer: str) -> str: return self.client.complete( "Is this answer correct?\n" f"Question: {question}\n" f"Answer: {answer}\n" "Return PASS or FAIL." ) client = ChatClient(model="gpt-4o", family="gpt") grader = Grader(client) answer = grader.extract_answer(apprentice_transcript) verdict = grader.judge_correctness(question, answer) ``` The object names imply a verifier, but there is no family boundary. The same client supplies the extraction and the judgment, and no assertion checks whether the verifier is independent from the generator. AutoGen's task-centric memory grader has this shape: a grader utility stores one `ChatCompletionClient` and uses it to extract the answer and judge correctness. Same-client grading is not wrong in every context. As a lightweight sample it is fine. It becomes a Cross-Family antipattern when the same-client verdict is treated as independent verification. The same failure family appears in other shapes. The zen pipeline uses same-family tiers for judge review, judge fix, and escalation. ChatArena shows that role-diverse agents can still leave the judge in the same family. DeepEval exposes an evaluation-model boundary, but a leaky default can leave callers with a same-family judge unless they choose otherwise. ### Pattern: uncovered verification-time instance No strict verification-time Cross-Family Pattern instance was found in the OSS bench surveyed for this catalog. The closest OSS evidence is DSPy's teacher/student split, which applies the same mechanic at optimization time rather than verification time. We mark this Pattern as uncovered rather than promoting an analogy-grade instance to canonical. DSPy still matters as supporting evidence. Its optimization code separates the model being shaped from the model proposing instructions, generating examples, or acting as teacher. That is the same independence move at a different leverage point: the optimization advice does not have to come from the model being optimized. When a strict verification-time instance is mined or self-mined, this card should be re-authored with the load-bearing assertion `verifier.family != generator.family` in the canonical block. ## Determinism Move Cross-Family constrains `same_family_bias` by requiring the verifier's family to be a configured value distinct from the generator's, recorded in the verdict so the boundary is auditable. It constrains `self_review_bias` by treating "different agent, same family" as a softer form of self-review rather than independent verification. The verifier may have a separate prompt and role, but it can still share the generator's blind spots. The determinism move is recorded family diversity at the verification boundary; if generator and verifier identities are not both in the verdict, the run is unauditable. ## Observable Signal Every Cross-Family report should include: * generator model and family; * verifier model and family; * family equality boolean (`generator_family == verifier_family`); * verdict; * verifier configuration source (`explicit_config`, `framework_default`, `module_singleton`); * Judge Harness fields if wrapped, such as consistency, perturbation, and calibration, or `harness: none` if not. A useful report names the model boundary: ```text generator_model: claude-sonnet-4-5 generator_family: claude verifier_model: gpt-4o verifier_family: gpt family_diverse: true verdict: pass verifier_config_source: explicit_config harness: none ``` ## Failure Modes Family inequality is a configuration fact; its effect on correlated error must be measured on the target workload, not assumed from the configuration. * **Same-Client Grading:** one client object is used for both generator and verifier. No assertion enforces family diversity. Construct distinct clients and assert at routing time. * **Leaky Default:** the framework exposes a configurable verifier slot, but the default resolves to the same family the application is built on. Refuse to construct the verifier when the family is unset. * **Role Diversity Mistaken for Verifier Diversity:** players, writers, managers, and researchers are diverse across families, but the judge defaults to one family. Identify the verification boundary separately from role assignment. * **Unrecorded Identities:** generation and verification use different families in practice, but the verdict does not name them. Stamp `generator_*` and `verifier_*` fields into every verdict. ## Use When Use this pattern when: * an LLM-based judge or evaluator is on the verification path; * the same provider serves both generation and verification by default; * the verification result gates promotion, training, deployment, or downstream automation; * optimization-time advice steers a model being shaped; * audit context requires verifier identity to be recorded. ## Do Not Use When Do not reach for Cross-Family when: * verification is fully executable, so an **Executable Analog** or **Comparator** can decide it without an LLM judge; * only one provider is reachable and provider diversity cannot be added; * the verification is low leverage and misattribution cost is negligible; * the boundary cannot carry verifier identity metadata. If only one family is reachable, label the result as informal and maximize executable checks. ## Evidence * **Verification Design Principle 7:** the design doc names Cross-Family Beats Self-Verification and frames same-family review as weaker than independent verification. * **Judge Reliability Harness update:** [the synthesis claim on LLM-judge reliability](https://github.com/verificationdesign/verificationdesign/blob/main/research/synthesis.md#llm-judge-reliability-can-vary-across-benchmarks-and-perturbations-even-when-the-judge-comes-from-a-different-model-family) records the caveat that judge reliability still needs perturbation, consistency, and calibration checks. * **[AutoGen](https://github.com/microsoft/autogen) task-centric memory grader:** the antipattern cleanup sweep records a same-client grader that extracts and judges answers with one `ChatCompletionClient`. * **zen same-family tiers:** the mining note records a provider-wide same-family pipeline where review, fix, and escalation stay inside the chosen family. * **[ChatArena](https://github.com/Farama-Foundation/chatarena) and [DeepEval](https://github.com/confident-ai/deepeval):** the cross-family sweep records role-diverse or configurable evaluation surfaces that do not, by themselves, enforce family diversity. * **[DSPy](https://github.com/stanfordnlp/dspy) teacher/student optimization:** the cross-family sweep records a partial Pattern instance at optimization time, where the teacher or proposer model is separate from the student model being shaped. ## Related Patterns * **Judge Harness:** cross-family is necessary, not sufficient; the verifier still needs perturbation, repetition, calibration, and reporting around it. * **Constitution:** a shared rubric makes cross-family verdicts comparable across runs. * **Admissibility Gate:** same-family verification with a confirmatory prompt doubles the antipattern; adversarial framing on a cross-family judge is the additive form. * **Blind Oracle:** Cross-Family addresses which model verifies; Blind Oracle addresses what the verifier is allowed to see. * **Adversary:** an adversary role drawn from the same family as the proposer reproduces the Cross-Family antipattern at the role boundary. --- Canonical: https://verificationdesign.com/patterns/orchestration/adversary/ Source: ai-design-patterns/cards/adversary.md # Adversary *(Orchestration Pattern)* ## Name **Adversary** Also known as: Critic, Red-Team Role, Negative Channel. ## Intent Assign a structurally separate role whose only job is to find failures in another role's output, and require that role to emit a negative channel the orchestrator can inspect. ## Problem A proposer produces work. The same proposer is then asked to "also list weaknesses," or a critic role is placed beside the proposer but given the same context and an optional feedback prompt. That can look like verification: * the transcript contains a message from a role named `critic`; * the prompt says "be critical"; * the critic gives suggestions before approval; * the workflow can continue when the critic says the work looks acceptable. The boundary still collapses when the proposer grades its own work or when the negative channel is optional. A critic that can return an empty feedback string without recording "no defect found" has not performed an adversarial pass. It has only added a chance for the model to rationalize the draft. `verification_design.md` Principle 1 rejects self-review as a verification signal. Principle 7 names the stronger form: cross-family verification beats self-verification. Adversary is the single-role orchestration primitive underneath those principles. It makes who critiques whom a runtime fact, not a tone instruction. ## Forces * **Separate role vs. single-agent self-critique.** A second role costs tokens and routing complexity; self-critique is cheaper but preserves the same blind spot. * **Shared context vs. blind critique.** Full context is easy to pass, but it can contaminate the critic. A Blind Oracle or Cross-Family verifier may be needed for stronger independence. * **Mandatory negative channel vs. optional feedback.** Optional feedback collapses to "looks good." A mandatory negative channel must either list defects or record an explicit no-defect verdict. * **Single critic vs. panel.** One adversary is the primitive. Multi-round disagreement belongs to Debate. * **Same family vs. cross-family.** A same-family adversary can still share latent priors with the proposer. Cross-Family strengthens the role boundary. ## Solution Make adversarial assignment explicit in code. The orchestrator names a proposer, names a critic, rejects `critic_id == proposer_id`, and requires a structured critique artifact. The artifact must contain: * proposer identity; * critic identity; * a reference to the artifact being critiqued; * weaknesses, including risks or rejected assumptions; * suggestions or next action; * a score or verdict, where this card's sample score is 0-100 and higher means more severe concern; * an explicit `no_defect_found` verdict if no weakness is found. The role label is not enough. The load-bearing structure is identity separation plus a required negative channel. ## Mechanism 1. **Assign role identities.** Give each proposal an `author_id`, and give each critique a distinct `critic_id`. 2. **Reject self-critique.** The orchestrator refuses to route a proposal back to its author as the adversary. 3. **Run the critic under a findings schema.** The critic returns structured weaknesses, suggestions, score, and verdict. 4. **Require a negative channel.** A `defects_found` verdict must include at least one weakness, and a `no_defect_found` verdict must include zero weaknesses. 5. **Route findings onward.** Findings gate release, route to Backpressure, or escalate. Routing is outside this card; the adversary only creates the failure signal. ## Pattern / Antipattern The same task: evaluate a proposal before it can advance. The antipattern side is intentionally uncovered in this pass. The pattern side shows the minimal identity and negative-channel assertions a verifier can inspect. ### Antipattern: uncovered confirmatory-critic instance No strict Adversary antipattern was promoted from the OSS bench surveyed for this catalog. The natural failure shape is a confirmatory critic: a role named `critic` or `adversary` that shares the proposer's context, asks for constructive feedback, and can approve without recording weaknesses or an explicit no-defect verdict. That shape is already covered by **Admissibility Gate**. A same-family critic that is treated as independent evidence is already covered by **Cross-Family**. This card keeps the Antipattern instance empty rather than inventing a second copy of those failures. When a strict instance is mined, re-author this section around the assertion `critic_id == proposal.author_id or negative_channel_present is False`. ### Pattern: separate critic with mandatory findings The structured implementation refuses self-critique, binds the returned artifact to the routed proposal and critic, and validates that the verdict agrees with the negative channel. ```python from dataclasses import dataclass from typing import Literal Verdict = Literal["defects_found", "no_defect_found"] @dataclass(frozen=True) class Proposal: proposal_id: str author_id: str content: str @dataclass(frozen=True) class Critique: proposal_id: str proposer_id: str critic_id: str weaknesses: tuple[str, ...] suggestions: tuple[str, ...] score: int verdict: Verdict def require_adversary(proposal: Proposal, critic_id: str) -> None: if critic_id == proposal.author_id: raise ValueError("critic must be distinct from proposer") def require_bound_critique( proposal: Proposal, requested_critic_id: str, critique: Critique ) -> None: if critique.proposal_id != proposal.proposal_id: raise ValueError("critique proposal_id must match proposal") if critique.proposer_id != proposal.author_id: raise ValueError("critique proposer_id must match proposal author") if critique.critic_id == proposal.author_id: raise ValueError("returned critic must be distinct from proposer") if critique.critic_id != requested_critic_id: raise ValueError("critique critic_id must match routed critic") def require_negative_channel(critique: Critique) -> None: if critique.verdict == "defects_found" and not critique.weaknesses: raise ValueError("defects_found requires at least one weakness") if critique.verdict == "no_defect_found" and critique.weaknesses: raise ValueError("no_defect_found requires zero weaknesses") def run_adversary(proposal: Proposal, critic_id: str, critic_fn) -> Critique: require_adversary(proposal, critic_id) critique = critic_fn(proposal=proposal, critic_id=critic_id) require_bound_critique(proposal, critic_id, critique) require_negative_channel(critique) return critique def build_adversary_report(proposal: Proposal, critique: Critique) -> dict[str, object]: return { "proposal_id": proposal.proposal_id, "proposer_id": critique.proposer_id, "critic_id": critique.critic_id, "critic_distinct_from_proposer": critique.critic_id != proposal.author_id, "negative_channel_present": bool(critique.weaknesses) or critique.verdict == "no_defect_found", "weakness_count": len(critique.weaknesses), "critique_score": critique.score, "verdict": critique.verdict, } def render_adversary_report(report: dict[str, object]) -> str: lines = [ f"proposal_id: {report['proposal_id']}", f"proposer_id: {report['proposer_id']}", f"critic_id: {report['critic_id']}", "critic_distinct_from_proposer: " f"{str(report['critic_distinct_from_proposer']).lower()}", "negative_channel_present: " f"{str(report['negative_channel_present']).lower()}", f"weakness_count: {report['weakness_count']}", f"critique_score: {report['critique_score']}", f"verdict: {report['verdict']}", ] return "\n".join(lines) proposal = Proposal( proposal_id="p-017", author_id="planner", content="Ship the migration without a rollback check.", ) def critic_fn(proposal: Proposal, critic_id: str) -> Critique: return Critique( proposal_id=proposal.proposal_id, proposer_id=proposal.author_id, critic_id=critic_id, weaknesses=("No rollback check is defined.",), suggestions=("Add a rollback verification gate before release.",), score=42, verdict="defects_found", ) critique = run_adversary(proposal, critic_id="critic", critic_fn=critic_fn) report = build_adversary_report(proposal, critique) def expect_value_error(expected_message: str, action) -> None: try: action() except ValueError as exc: assert str(exc) == expected_message else: raise AssertionError(f"expected ValueError: {expected_message}") expect_value_error( "critic must be distinct from proposer", lambda: run_adversary(proposal, critic_id="planner", critic_fn=critic_fn), ) def bad_critique(**overrides) -> Critique: values = { "proposal_id": "p-999", "proposer_id": "other-planner", "critic_id": "critic-shadow", "weaknesses": ("A different flaw is reported.",), "suggestions": ("Route the proposal through the matching reviewer.",), "score": 90, "verdict": "defects_found", } values.update(overrides) return Critique(**values) binding_cases = ( ( "critique proposal_id must match proposal", bad_critique( proposer_id=proposal.author_id, critic_id="critic", ), ), ( "critique proposer_id must match proposal author", bad_critique( proposal_id=proposal.proposal_id, critic_id="critic", ), ), ( "critique critic_id must match routed critic", bad_critique( proposal_id=proposal.proposal_id, proposer_id=proposal.author_id, ), ), ( "returned critic must be distinct from proposer", bad_critique( proposal_id=proposal.proposal_id, proposer_id=proposal.author_id, critic_id=proposal.author_id, ), ), ) for expected_message, returned_critique in binding_cases: expect_value_error( expected_message, lambda returned_critique=returned_critique: run_adversary( proposal, critic_id="critic", critic_fn=lambda **_: returned_critique, ), ) expect_value_error( "no_defect_found requires zero weaknesses", lambda: run_adversary( proposal, critic_id="critic", critic_fn=lambda **_: Critique( proposal_id=proposal.proposal_id, proposer_id=proposal.author_id, critic_id="critic", weaknesses=("Rollback is still unchecked.",), suggestions=("Add the missing rollback gate.",), score=65, verdict="no_defect_found", ), ), ) assert critique.critic_id != proposal.author_id assert report["negative_channel_present"] is True assert report["critic_distinct_from_proposer"] is True assert render_adversary_report(report) == """proposal_id: p-017 proposer_id: planner critic_id: critic critic_distinct_from_proposer: true negative_channel_present: true weakness_count: 1 critique_score: 42 verdict: defects_found""" ``` AutoGPT's `multi_agent_debate.py` has this shape as a legacy v1 instance. Its critique artifact records `critic_id`, `target_agent_id`, weaknesses, suggestions, and score. Its critique phase skips self-critique by skipping `j == i`, so a proposal owner does not critique itself. A silent skip has different failure semantics from the sample above: it can leave a proposal uncritiqued unless the surrounding debate loop accounts for missing pairings. That same AutoGPT instance also belongs on Debate. The critique artifact and self-critique exclusion are the Adversary mechanic; the bounded multi-round exchange and consensus behavior are the Debate mechanic. AutoGen's writer and critic example in the migration guide is a partial instance. The critic is a named role in a `RoundRobinGroupChat`, and `TextMentionTermination("APPROVE")` makes explicit approval the release condition. That shape is also Backpressure because unresolved critic feedback keeps the loop running. ## Determinism Move Adversary constrains `self_review_bias` by making the proposer unable to satisfy the adversarial step alone. The critic identity is external to the proposal, and the assertion rejects `critic_id == proposer_id`. A same-family critic can still share the proposer's blind spots. Adversary by itself does not constrain `same_family_bias`; the recorded role boundary is where Cross-Family can attach its enforced family check. The determinism move is making the negative channel mandatory and the critic's identity external. ## Observable Signal Every Adversary report should include: * proposal id; * proposer id; * critic id; * critic distinct from proposer boolean; * negative-channel present boolean; * weakness count; * critique score, on a 0-100 concern scale where higher means more severe; * verdict. A useful report makes the role boundary visible: ```text proposal_id: p-017 proposer_id: planner critic_id: critic critic_distinct_from_proposer: true negative_channel_present: true weakness_count: 1 critique_score: 42 verdict: defects_found ``` Downstream orchestration, such as Backpressure or Escalation Chain, records routing decisions separately. ## Failure Modes * **Confirmatory Critic:** the role is named critic, but the prompt asks for constructive feedback and approval. Use Admissibility Gate so the critic must search for failure before approval. * **Self-Critique:** the critic and proposer are the same role or model call. Assert identity separation before routing the critique. * **Optional Negative Channel:** the critic can return empty feedback with no recorded `no_defect_found` verdict. Reject empty critiques unless the no-defect verdict is explicit. * **Toothless Adversary:** findings are produced but never gate, revise, or escalate. Connect the report to Backpressure or Escalation Chain. ## Use When Use this pattern when: * a single proposer's blind spots are costly; * the workflow can afford a second role; * the system needs an explicit failure-search step before release; * the critique should create a routable artifact, not only prose; * later Backpressure, Escalation Chain, or Debate steps need a negative signal. ## Do Not Use When Do not reach for Adversary when: * the task is trivial and a second role would add process noise; * the critic would be the same model, same prompt context, and same family, with no recorded independence; * an Executable Analog or Comparator can decide the property without an LLM critic; * the desired structure is multi-round disagreement among several roles. Use Debate for that. If only a same-family critic is available, label the result as a weak adversarial pass and maximize executable checks around it. ## Evidence * **Verification Design Principles 1 and 7:** the design doc rejects self-review as a verification signal and frames independent verification as stronger than same-family review. * **[AutoGPT](https://github.com/Significant-Gravitas/AutoGPT) multi-agent debate:** the orchestration sweep records a direct Adversary instance: `AgentCritique` names critic and target identities, records weaknesses and suggestions, and skips self-critique in the critique phase. * **AutoGPT legacy caveat:** the same evidence comes from AutoGPT's legacy v1 codebase, so it is treated as a historical implementation, not a current framework recommendation. * **[AutoGen](https://github.com/microsoft/autogen) writer and critic migration guide:** the orchestration sweep records a partial instance where a critic role must emit `APPROVE` before the writer/critic loop terminates. * **No promoted antipattern:** the orchestration sweep did not promote a strict Adversary antipattern; this card cross-references Admissibility Gate and Cross-Family instead of inventing one. ## Related Patterns * **Admissibility Gate**: defines the default-no and admissibility logic an adversary applies to each finding. * **Cross-Family:** strengthens the adversary by making the critic come from a different model family. * **Debate:** generalizes Adversary into multi-round, multi-critic disagreement. * **Escalation Chain:** receives unresolved adversary findings when the critic cannot safely approve. * **Backpressure:** routes adversary findings back to the proposer for revision. --- Canonical: https://verificationdesign.com/patterns/orchestration/debate/ Source: ai-design-patterns/cards/debate.md # Debate *(Orchestration Pattern)* ## Name **Debate** Also known as: Structured Disagreement, Multi-Agent Debate, Roundtable. ## Intent Run bounded disagreement among multiple roles before a decision, with turn order, round count, phase, and consensus threshold held in orchestration state instead of model discretion. ## Problem Debate is easy to imitate and hard to get as structure. The theatrical version looks busy: * one model is asked to play several debaters; * every role receives the same prompt and the same transcript; * a model decides who should speak next; * "consensus" appears when the transcript sounds settled; * the discussion either converges immediately or keeps going until the context budget runs out. That is not a debate mechanism. It is a role-play transcript. The system cannot tell whether disagreement was real, whether every participant got a turn, whether the decision followed a threshold, or whether the model simply declared agreement. `verification_design.md` Principle 8 names the useful move: simulate debate before concluding. For systems work, the simulation has to be orchestrated. The useful artifact is not "several messages happened." It is a bounded record of who spoke, in which phase, under which stopping rule, and how the decision was counted. ## Forces * **Single critic vs. multi-round disagreement.** Adversary is cheaper and often enough. Debate pays for several roles and rounds when the decision benefits from multiple positions. * **Independent positions vs. shared-transcript contamination.** A shared transcript lets debaters respond to one another, but it can also pull them toward premature agreement. * **Round cap vs. convergence detection.** A fixed cap is auditable. Model-detected convergence is flexible but can hide the stopping rule. * **Consensus threshold vs. unanimity vs. majority.** The decision rule should be named before the debate starts. * **Deterministic speaker schedule vs. model-chosen speaker.** A schedule is less expressive than a moderator model, but it makes turn taking inspectable. ## Solution Externalize the debate state. The orchestrator registers debater identities, selects the next speaker by a deterministic schedule, records each position under a schema, and applies a named decision rule. Consensus is not whatever the final model says. It is a counted condition, such as a threshold over votes, or a bounded failure to reach consensus after `max_rounds`. The model may generate arguments. Code owns: * participant ids; * phase; * current speaker; * round count; * maximum rounds; * consensus threshold; * vote tally; * final decision rule. ## Mechanism 1. **Register debaters.** Store participant ids and the order in which they speak. 2. **Initialize debate state.** Set phase, round count, maximum rounds, consensus threshold, and decision rule. 3. **Select speakers by schedule.** The orchestrator chooses the next speaker, not a model. 4. **Collect structured positions.** Each turn records speaker id, round, stance, rationale, and vote. 5. **Apply the stopping rule.** Stop when the threshold is met by distinct debaters' current votes or when the round cap is exhausted, and emit the vote tally and decision. The threshold check runs after every phase, so a debate can stop before critique or revision when the proposal votes already meet the rule. If the cap is exhausted without consensus, route the unresolved outcome to escalation. ## Pattern / Antipattern The same task: decide whether to release a risky migration after structured disagreement. The antipattern side is intentionally uncovered in this pass. The pattern side shows a minimal bounded loop where speaker selection and consensus are code-level facts. ### Antipattern: uncovered theatrical-debate instance No strict Debate antipattern was promoted from the OSS bench surveyed for this catalog. The natural failure shape is theatrical debate: role labels without independent positions, no round cap, no speaker schedule, and a model-declared consensus. That overlaps with **Adversary** when the "critic" is only a label, and with **Cross-Family** when all debaters share the same family and priors. This card keeps the Antipattern instance empty rather than fabricating a code example. When a strict instance is mined, re-author this section around observable model-chosen speaking or a missing counted threshold. ### Pattern: bounded round-robin debate The structured implementation stores turn order and stopping rules outside the model. ```python from collections import Counter from dataclasses import dataclass from typing import Literal Phase = Literal["proposal", "critique", "revision", "consensus"] DecisionRule = Literal["threshold_vote", "max_rounds_exhausted"] Vote = Literal["release", "revise", "escalate"] @dataclass(frozen=True) class Turn: round_index: int speaker_id: str phase: Phase stance: str rationale: str vote: Vote @dataclass(frozen=True) class TurnContent: stance: str rationale: str vote: Vote @dataclass(frozen=True) class DebateResult: debater_ids: tuple[str, ...] max_rounds: int rounds_run: int consensus_threshold: int speaker_schedule: tuple[str, ...] phase_sequence: tuple[Phase, ...] vote_tally: dict[Vote, int] decision: Vote decision_rule: DecisionRule speaker_selected_by: Literal["schedule"] transcript: tuple[Turn, ...] def to_report(self) -> str: debaters = ", ".join(self.debater_ids) phases = ", ".join(self.phase_sequence) tally = ", ".join( f"{vote}: {count}" for vote, count in self.vote_tally.items() ) schedule = ", ".join(self.speaker_schedule) return "\n".join( [ f"debater_ids: [{debaters}]", f"rounds_run: {self.rounds_run}", f"max_rounds: {self.max_rounds}", f"phase_sequence: [{phases}]", f"consensus_threshold: {self.consensus_threshold}", f"vote_tally: {{{tally}}}", f"decision: {self.decision}", f"decision_rule: {self.decision_rule}", f"speaker_schedule: [{schedule}]", ] ) def latest_vote_tally(transcript: list[Turn]) -> Counter[Vote]: latest_vote_by_speaker = { turn.speaker_id: turn.vote for turn in transcript } return Counter(latest_vote_by_speaker.values()) def run_debate( debater_ids: tuple[str, ...], max_rounds: int, consensus_threshold: int, collect_turn, ) -> DebateResult: if not debater_ids: raise ValueError("debater_ids must not be empty") if max_rounds < 1: raise ValueError("max_rounds must be at least 1") transcript: list[Turn] = [] phases: list[Phase] = [] schedule: list[str] = [] decision: Vote | None = None decision_rule: DecisionRule = "max_rounds_exhausted" votes: Counter[Vote] = Counter() for round_index in range(1, max_rounds + 1): for phase in ("proposal", "critique", "revision", "consensus"): phases.append(phase) for speaker_id in debater_ids: schedule.append(speaker_id) content = collect_turn( round_index=round_index, speaker_id=speaker_id, phase=phase, ) turn = Turn( round_index=round_index, speaker_id=speaker_id, phase=phase, stance=content.stance, rationale=content.rationale, vote=content.vote, ) transcript.append(turn) votes = latest_vote_tally(transcript) winner, count = votes.most_common(1)[0] if count >= consensus_threshold: decision = winner decision_rule = "threshold_vote" break if decision is not None: break if decision is None: votes = latest_vote_tally(transcript) decision = "escalate" return DebateResult( debater_ids=debater_ids, max_rounds=max_rounds, rounds_run=transcript[-1].round_index, consensus_threshold=consensus_threshold, speaker_schedule=tuple(schedule), phase_sequence=tuple(phases), vote_tally=dict(votes), decision=decision, decision_rule=decision_rule, speaker_selected_by="schedule", transcript=tuple(transcript), ) def collect_turn(round_index: int, speaker_id: str, phase: Phase) -> TurnContent: vote_by_speaker = { "planner": "release", "critic": "revise", "operator": "revise", } return TurnContent( stance=f"{speaker_id} position during {phase}", rationale=f"{speaker_id} rationale during {phase}", vote=vote_by_speaker[speaker_id], ) result = run_debate( debater_ids=("planner", "critic", "operator"), max_rounds=2, consensus_threshold=2, collect_turn=collect_turn, ) assert result.decision == "revise" assert result.vote_tally == {"release": 1, "revise": 2} assert result.decision_rule == "threshold_vote" assert result.rounds_run == 1 assert tuple(turn.speaker_id for turn in result.transcript) == result.speaker_schedule assert result.to_report() == """debater_ids: [planner, critic, operator] rounds_run: 1 max_rounds: 2 phase_sequence: [proposal] consensus_threshold: 2 vote_tally: {release: 1, revise: 2} decision: revise decision_rule: threshold_vote speaker_schedule: [planner, critic, operator]""" def split_vote_turn(round_index: int, speaker_id: str, phase: Phase) -> TurnContent: vote_by_speaker = { "planner": "release", "critic": "revise", "operator": "escalate", } return TurnContent( stance=f"{speaker_id} split position during {phase}", rationale=f"{speaker_id} split rationale during {phase}", vote=vote_by_speaker[speaker_id], ) split_result = run_debate( debater_ids=("planner", "critic", "operator"), max_rounds=2, consensus_threshold=2, collect_turn=split_vote_turn, ) assert split_result.decision_rule == "max_rounds_exhausted" assert split_result.decision == "escalate" assert split_result.rounds_run == split_result.max_rounds assert split_result.vote_tally == {"release": 1, "revise": 1, "escalate": 1} ``` AutoGPT's legacy `multi_agent_debate.py` has this shape as a v1 instance. It names proposal, critique, revision, consensus, and execution phases; stores proposal and critique artifacts; exposes debater count, round count, consensus threshold, and voting mode; and moves through bounded rounds before consensus. AutoGen's `RoundRobinGroupChat` is the speaker-schedule instance. Its manager persists the message thread, current turn, and next-speaker index, then selects exactly one participant per turn by round-robin order. Debate is orchestration state, not a model choosing who should speak next. ## Determinism Move Debate constrains `criteria_drift` by putting the decision rule in code. Consensus means a named threshold over current votes by distinct debaters, a vote tally, or exhausted round budget, not the model's changing sense that the group now agrees. Debate also creates an inspectable record of who argued what and whether one role dominated. By itself, it does not constrain `self_review_bias`; use Cross-Family and Adversary when the system must enforce producer-verifier separation. The determinism move is putting the stopping rule and turn order in code, so consensus is a counted condition, not a vibe. ## Observable Signal Every Debate report should include: * debater ids; * rounds run; * maximum rounds; * phase sequence; * consensus threshold; * vote tally; * decision; * decision rule; * speaker schedule. A useful report shows the counted decision: ```text debater_ids: [planner, critic, operator] rounds_run: 1 max_rounds: 2 phase_sequence: [proposal] consensus_threshold: 2 vote_tally: {release: 1, revise: 2} decision: revise decision_rule: threshold_vote speaker_schedule: [planner, critic, operator] ``` ## Failure Modes * **Theatrical Debate:** role labels exist, but every debater shares the same prompt, same model family, and same priors. Use Cross-Family when the disagreement needs real diversity. * **Unbounded Debate:** there is no round cap, and a model decides when the group has converged. Set `max_rounds` and record why the loop stopped. * **Model-Chosen Speaker:** the next speaker is chosen by a model rather than a schedule. Record the speaker-selection rule or use a round-robin order. * **Phantom Consensus:** consensus is declared without a threshold or vote count. Require a named decision rule and a tally. ## Use When Use this pattern when: * a high-stakes decision benefits from several perspectives; * a single Adversary pass is too narrow; * the system can afford multiple roles and bounded rounds; * the decision needs an auditable vote, threshold, or turn record; * unresolved disagreement should route to Escalation Chain or Backpressure. ## Do Not Use When Do not reach for Debate when: * a single critic is enough. Use Adversary; * an Executable Analog or Comparator can decide the property directly; * token, latency, or provider cost cannot support multiple turns; * all debaters are the same model, same prompt, and same family, with no recorded diversity. If the debate cannot create independent positions or a counted stopping rule, do not call it debate. ## Evidence * **Verification Design Principle 8:** the design doc names Simulate Debate and frames debate as a way to force counterargument before conclusion. * **[AutoGPT](https://github.com/Significant-Gravitas/AutoGPT) multi-agent debate:** the orchestration sweep records a direct Debate instance with explicit phases, separate proposal and critique artifacts, configurable debater count, round count, consensus threshold, and voting mode. * **AutoGPT legacy caveat:** the same implementation lives in AutoGPT's legacy classic tree, so it is treated as a legacy v1 implementation rather than a current framework recommendation. * **[AutoGen](https://github.com/microsoft/autogen) RoundRobinGroupChat:** the orchestration sweep records a direct speaker-schedule instance: message thread, current turn, next-speaker index, one selected participant per turn, max turns, and termination conditions. * **No promoted antipattern:** the orchestration sweep did not promote a strict Debate antipattern; this card cross-references Adversary and Cross-Family rather than inventing one. ## Related Patterns * **Adversary:** Debate generalizes the single-critic primitive into multi-round, multi-role disagreement. * **Cross-Family:** makes debater diversity stronger by separating model families. * **Escalation Chain:** receives unresolved debate outcomes when no safe consensus is reached. * **Backpressure:** routes debate findings back to revision. * **Admissibility Gate:** defines the default-no posture each debater can apply to other positions. --- Canonical: https://verificationdesign.com/patterns/orchestration/escalation-chain/ Source: ai-design-patterns/cards/escalation-chain.md # Escalation Chain *(Orchestration Pattern)* ## Name **Escalation Chain** Also known as: Handoff, Tiered Routing, Manager Delegation. ## Intent Route work to a higher-authority or different-capability handler through a typed, validated handoff, so the next owner is code-level state instead of a model's memory of who to call. ## Problem Escalation is often written as a prompt convention: * "If you cannot handle this, ask the manager." * "Escalate low-confidence answers." * "Call a senior reviewer when uncertain." * "Route failures to another agent." Those instructions name a desired behavior, not a routing mechanism. The model still decides whether escalation is needed, which target to call, and whether the target is different from itself. It may retry the same failing handler, route to an unknown role, or send the work to the same model family that produced the failure. No typed handoff means no inspectable target. No validated target means no guarantee the next handler exists. No depth cap means escalation can loop. `verification_design.md` Principle 2 names the core independence rule: generation and verification should be separated so the verifier does not copy the same errors. Escalation Chain extends that rule to routing by analogy. A failed or low-confidence handler should transfer control to a validated next handler, not ask the same model to remember an escalation instruction. Principle 7 adds the same-family caveat: escalation within the same family is weaker than escalation across a real family or capability boundary. ## Forces * **Prompt convention vs. typed handoff.** A sentence can be ignored or misremembered. A handoff object can be validated. * **Model-chosen target vs. orchestrator-enforced target.** Letting the model choose is flexible, but it makes the next owner unauditable unless the target is checked against known participants. * **Same-family escalation vs. independent escalation.** Routing to a sibling role in the same family is cheap, but it may preserve the same blind spot. * **Bounded depth vs. unbounded loops.** A chain needs a depth cap so failure does not bounce forever. * **Manager process vs. ad hoc delegation.** A manager adds ceremony, but it centralizes authority and delegation tooling. ## Solution Represent escalation as a typed message or configured manager process. The orchestrator defines known handlers, allowed targets, maximum depth, and the authority that may route work. When a handler cannot safely proceed, it emits a handoff artifact naming: * source handler; * target handler; * reason; * payload or work item. The orchestrator validates the source and target against registered participants, checks that the source is the current handler, rejects self-escalation, enforces the allowed-target policy, computes depth from the recorded path, enforces the depth cap, and records the path. Retries are a Backpressure loop, not a self-handoff through this escalation boundary. Escalation authority belongs in code. The model may propose a handoff. The orchestrator decides whether the handoff is valid. ## Mechanism 1. **Define handlers and tiers.** Register handler ids, capabilities, families, and allowed handoff targets. 2. **Define a handoff schema.** The schema carries source, target, reason, and payload. The orchestrator derives depth from the recorded path. 3. **Validate the route.** Reject unknown sources, unknown targets, source impersonation, self-escalation, disallowed transitions, and over-depth paths. 4. **Route through the orchestrator.** The orchestrator sets the next handler from the validated handoff, not from free prose. 5. **Record the escalation path.** Store source, target, reason, computed depth, max depth, performed checks, process mode, and path. ## Pattern / Antipattern The same task: route a failed verification to a stronger handler. The antipattern side is intentionally uncovered in this pass. The pattern side shows a typed handoff whose target is validated before control moves. ### Antipattern: uncovered prompt-convention escalation No strict Escalation Chain antipattern was promoted from the OSS bench surveyed for this catalog. The natural failure shape is prompt-convention escalation: a model is told to ask a manager when stuck, but the routing target is not typed, not validated, and can loop back to the same handler or same family. Same-family escalation is already covered by **Cross-Family**. Label-only routing roles overlap with **Adversary** when the role name implies independence that the code does not enforce. This card keeps the Antipattern instance empty rather than inventing one. When a strict instance is mined, re-author this section around the assertion `handoff.target not in participants or handoff.target == handoff.source`. ### Pattern: typed handoff with validated target The structured implementation makes the escalation target a validated fact before routing. ```python from dataclasses import dataclass @dataclass(frozen=True) class Handler: handler_id: str family: str tier: int @dataclass(frozen=True) class Handoff: source: str target: str reason: str payload: str @dataclass(frozen=True) class EscalationResult: source_handler: str target_handler: str escalation_reason: str escalation_depth: int max_depth: int checks_performed: tuple[str, ...] process_mode: str next_handler: Handler path: tuple[str, ...] def to_report(self) -> dict[str, object]: return { "source_handler": self.source_handler, "target_handler": self.target_handler, "escalation_reason": self.escalation_reason, "escalation_depth": self.escalation_depth, "max_depth": self.max_depth, "checks_performed": self.checks_performed, "process_mode": self.process_mode, "path": self.path, } def route_handoff( handoff: Handoff, participants: dict[str, Handler], allowed_targets: dict[str, set[str]], max_depth: int, path: tuple[str, ...], ) -> EscalationResult: checks = ( "unknown_target", "source_identity", "allowed_transition", "self_escalation", "depth_cap", ) if handoff.source not in participants: raise ValueError(f"unknown escalation source: {handoff.source}") if handoff.target not in participants: raise ValueError(f"unknown escalation target: {handoff.target}") if not path or path[-1] != handoff.source: raise ValueError("handoff source is not the current handler") if handoff.target == handoff.source: raise ValueError("self-escalation is not an escalation") if handoff.target not in allowed_targets.get(handoff.source, set()): raise ValueError("target is not allowed for this source") depth = len(path) if depth > max_depth: raise ValueError("escalation depth exceeded") return EscalationResult( source_handler=handoff.source, target_handler=handoff.target, escalation_reason=handoff.reason, escalation_depth=depth, max_depth=max_depth, checks_performed=checks, process_mode="typed_handoff", next_handler=participants[handoff.target], path=path + (handoff.target,), ) participants = { "writer": Handler(handler_id="writer", family="gpt", tier=1), "reviewer": Handler(handler_id="reviewer", family="claude", tier=2), "human": Handler(handler_id="human", family="human", tier=3), } allowed_targets = { "writer": {"reviewer"}, "reviewer": {"human"}, "human": set(), } handoff = Handoff( source="writer", target="reviewer", reason="low confidence after adversary finding", payload="migration plan requires independent review", ) result = route_handoff( handoff=handoff, participants=participants, allowed_targets=allowed_targets, max_depth=2, path=("writer",), ) expected_report = { "source_handler": "writer", "target_handler": "reviewer", "escalation_reason": "low confidence after adversary finding", "escalation_depth": 1, "max_depth": 2, "checks_performed": ( "unknown_target", "source_identity", "allowed_transition", "self_escalation", "depth_cap", ), "process_mode": "typed_handoff", "path": ("writer", "reviewer"), } def assert_rejected(bad_handoff: Handoff, bad_path: tuple[str, ...]) -> None: try: route_handoff( handoff=bad_handoff, participants=participants, allowed_targets=allowed_targets, max_depth=2, path=bad_path, ) except ValueError: return raise AssertionError("invalid handoff was accepted") assert_rejected( Handoff( source="writer", target="ghost", reason="unknown target", payload="work item", ), ("writer",), ) assert_rejected( Handoff( source="writer", target="writer", reason="self-target", payload="work item", ), ("writer",), ) assert_rejected( Handoff( source="reviewer", target="human", reason="wrong source", payload="work item", ), ("writer",), ) assert_rejected( Handoff( source="reviewer", target="writer", reason="disallowed transition", payload="work item", ), ("writer", "reviewer"), ) assert_rejected( Handoff( source="reviewer", target="human", reason="over depth", payload="work item", ), ("writer", "human", "reviewer"), ) assert result.next_handler.handler_id == handoff.target assert result.to_report() == expected_report ``` AutoGen's `Swarm` team has this typed-handoff shape. It selects the next speaker only from handoff messages, validates that handoff targets are participants, persists the current speaker, and otherwise keeps the current speaker when no handoff occurs. CrewAI's hierarchical process is the manager-process sibling. Its hierarchical mode requires a manager LLM or manager agent, routes hierarchical execution through a manager, and equips the manager with delegation tools. That is escalation as configured process mode rather than prompt convention. ## Determinism Move Escalation Chain constrains `tool_boundary_ambiguity` by making transfer of control happen at a typed boundary. The handoff target is validated against participants before the next handler runs, rather than being held in model memory. Escalating to a same-family sibling can still be useful for load or specialization. Escalation Chain by itself does not constrain `same_family_bias`; that requires Cross-Family's enforced family boundary at the verifier seam. The determinism move is making the routing target a validated fact in code, not a model's recollection of who to call. ## Observable Signal Every Escalation Chain report should include: * source handler; * target handler; * escalation reason; * escalation depth; * maximum depth; * checks performed; * process mode, such as `typed_handoff` or `manager_process`; * manager-present boolean for manager-process reports; * path so far. A useful report names the handoff: ```text source_handler: writer target_handler: reviewer escalation_reason: low confidence after adversary finding escalation_depth: 1 max_depth: 2 checks_performed: [unknown_target, source_identity, allowed_transition, self_escalation, depth_cap] process_mode: typed_handoff path: [writer, reviewer] ``` ## Failure Modes * **Prompt-Convention Escalation:** the prompt says "ask the manager," but no typed target is emitted or validated. Add a handoff object and reject unknown targets. * **Self-Escalation:** the target is the source handler. Reject self-handoffs at this boundary. Retries belong in Backpressure, not in Escalation Chain. * **Same-Family Escalation:** work routes to a role that shares the same model family and blind spot. Use Cross-Family when independence matters. * **Unbounded Escalation:** the chain has no depth cap and loops among handlers. Record depth and enforce `max_depth`. * **Unvalidated Target:** the handoff names a handler that is not registered. Reject before routing. ## Use When Use this pattern when: * failures or low-confidence decisions should route to a more capable handler; * the system has tiers, specialists, managers, or humans; * the next handler must be auditable; * unresolved Adversary or Debate findings need a stronger route; * same-family retry would be insufficient evidence of independence. ## Do Not Use When Do not reach for Escalation Chain when: * a flat single-handler design is enough; * an Executable Analog or Comparator can decide the property without routing; * the proposed target is the same model, same prompt, and same family; * there is no owner capable of handling the escalated work; * the chain cannot enforce a maximum depth. If escalation cannot name and validate a different next owner, label it as retry, not escalation. ## Evidence * **Verification Design Principle 2:** the design doc names independence between generation and verification; Escalation Chain applies that independence rule to routing by analogy after failure or low confidence. * **Verification Design Principle 7:** the same design doc frames cross-family verification as stronger than self-verification, which supplies the same-family caveat for escalation targets. * **[AutoGen](https://github.com/microsoft/autogen) Swarm:** the orchestration sweep records a direct typed-handoff instance: next speaker is selected from handoff messages, targets are validated against participants, and current speaker is persisted. * **[CrewAI](https://github.com/crewAIInc/crewAI) hierarchical process:** the orchestration sweep records a direct manager-process instance: hierarchical mode requires a manager, routes execution through that manager, and supplies delegation tools. * **No promoted antipattern:** the orchestration sweep did not promote a strict Escalation Chain antipattern; this card cross-references Cross-Family and Adversary rather than inventing one. ## Related Patterns * **Backpressure:** routes failure back upstream for revision; Escalation Chain routes failure up or out to a different handler. * **Cross-Family:** makes escalation independent when the target comes from a different model family. * **Debate:** unresolved debate can escalate when no safe consensus is reached. * **Adversary:** adversary findings can trigger escalation. * **Causal Tag:** tags escalation events so the routing path is auditable across logs. --- Canonical: https://verificationdesign.com/patterns/orchestration/backpressure/ Source: ai-design-patterns/cards/backpressure.md # Backpressure *(Orchestration Pattern)* ## Name **Backpressure** Also known as: Retry-with-Feedback, Rerun Context, Revision Loop. ## Intent When a downstream check fails, route the failure back upstream as structured rerun context within a bounded retry budget, instead of swallowing the failure or retrying blindly. ## Problem A downstream verifier fails. The system has several easy ways to lose that signal: * log the failure and continue downstream; * retry the same upstream call without telling it what failed; * pass back a raw error string the upstream step cannot act on; * keep retrying until a timeout or context budget stops the loop; * ask the model to "try again" with no record of why. That is upstream-unaware failure. The verifier found something, but the work step never receives a structured correction. The next attempt is sampled from the same prompt with the same blind spot. If it passes later, the system cannot tell whether the failure was fixed or merely disappeared. `verification_design.md` Principle 6 names the core rule: executable verification is king. Backpressure is the orchestration move after that executable or rubric-backed verifier fails. The failure should block progress, become structured context, and drive a bounded rerun. Principle 3 supplies the step-level framing: failures are most useful when they are attached to the step that produced them, not discovered only at the end. ## Forces * **Blind retry vs. feedback-carrying retry.** Blind retry may pass by chance. Feedback-carrying retry tells the upstream step what must change. * **Bounded budget vs. unbounded loop.** Retry is useful only when capped. * **Swallow failure vs. block progress.** Continuing downstream after failed validation turns a verifier into a logger. * **Structured rerun context vs. raw error.** A typed failure can be consumed. A traceback may only confuse the next step. * **Raise or escalate vs. silent give-up.** Exhaustion should become an explicit outcome, not a quiet stop. ## Solution Feed the failure back to the producer. The orchestrator runs the work step, runs the verifier, and if the verifier fails, formats the failure into rerun context. It invokes the work step again with that context and increments a retry counter. The loop stops only when the verifier passes or the budget is exhausted. On exhaustion, the loop returns an explicit exhaustion outcome for the caller to raise on or route to escalation. The model may revise the artifact. Code owns: * the pass/fail criterion; * retry count; * maximum retries; * failure context format; * per-attempt verdicts; * final outcome. ## Mechanism 1. **Run the work step.** Produce an artifact. 2. **Run the verifier.** The verifier returns a structured pass or failure with a reason. 3. **On failure, format rerun context.** Convert the failure into actionable context for the upstream step. 4. **Retry within a budget.** Reinvoke the upstream step with rerun context and increment attempts. 5. **Stop explicitly.** Proceed on pass; on exhaustion, record the outcome so the caller can raise or escalate. ## Pattern / Antipattern The same task: revise an output until a downstream validation passes or the retry budget is exhausted. The antipattern side is intentionally uncovered in this pass. The pattern side shows the failure becoming feedback within a counted budget. ### Antipattern: uncovered swallowed-failure instance No strict Backpressure antipattern was promoted from the OSS bench surveyed for this catalog. The natural failure shape is downstream-failure-swallowed, upstream-unaware: validation fails, but the system proceeds, logs only, or retries with no feedback context and no budget. That overlaps with **Admissibility Gate** when there is no failure-search rubric, so there is no structured failure to push back. This card keeps the Antipattern instance empty rather than inventing one. When a strict instance is mined, re-author this section around the assertion `attempts > 1 and rerun_context is None`, or `verifier_failed and proceeded_downstream is True`. ### Pattern: bounded retry with structured feedback The structured implementation formats verifier failure into rerun context and retries only within the budget. ```python from dataclasses import dataclass from typing import Literal Outcome = Literal["passed", "exhausted"] @dataclass(frozen=True) class Verification: passed: bool reason: str @dataclass(frozen=True) class Attempt: index: int artifact: str verdict: str rerun_context: str | None @dataclass(frozen=True) class BackpressureResult: attempts: int max_retries: int rerun_context_fed_back: bool last_failure_reason: str | None per_attempt_verdicts: tuple[str, ...] outcome: Outcome artifact: str attempt_log: tuple[Attempt, ...] def to_report(self) -> dict[str, object]: return { "attempts": self.attempts, "max_retries": self.max_retries, "rerun_context_fed_back": self.rerun_context_fed_back, "last_failure_reason": self.last_failure_reason, "per_attempt_verdicts": self.per_attempt_verdicts, "outcome": self.outcome, "artifact": self.artifact, } def format_rerun_context(check: Verification) -> str: return f"Validation failed: {check.reason}. Revise only this failure." def run_with_backpressure(work_fn, verify_fn, max_retries: int) -> BackpressureResult: rerun_context: str | None = None attempts: list[Attempt] = [] last_failure_reason: str | None = None for attempt_index in range(1, max_retries + 2): artifact = work_fn(rerun_context) check = verify_fn(artifact) verdict = "pass" if check.passed else "fail" attempts.append( Attempt( index=attempt_index, artifact=artifact, verdict=verdict, rerun_context=rerun_context, ) ) if check.passed: return BackpressureResult( attempts=attempt_index, max_retries=max_retries, rerun_context_fed_back=any(a.rerun_context for a in attempts), last_failure_reason=last_failure_reason, per_attempt_verdicts=tuple(a.verdict for a in attempts), outcome="passed", artifact=artifact, attempt_log=tuple(attempts), ) last_failure_reason = check.reason if attempt_index > max_retries: return BackpressureResult( attempts=attempt_index, max_retries=max_retries, rerun_context_fed_back=any(a.rerun_context for a in attempts), last_failure_reason=last_failure_reason, per_attempt_verdicts=tuple(a.verdict for a in attempts), outcome="exhausted", artifact=artifact, attempt_log=tuple(attempts), ) rerun_context = format_rerun_context(check) raise AssertionError("unreachable") def work_fn(rerun_context: str | None) -> str: if rerun_context is None: return "migration plan without rollback verification" if "rollback verification" in rerun_context: return "migration plan with rollback verification" return "migration plan without rollback verification" def verify_fn(artifact: str) -> Verification: if "with rollback verification" in artifact: return Verification(passed=True, reason="") return Verification(passed=False, reason="rollback verification is missing") result = run_with_backpressure(work_fn, verify_fn, max_retries=2) expected_report = { "attempts": 2, "max_retries": 2, "rerun_context_fed_back": True, "last_failure_reason": "rollback verification is missing", "per_attempt_verdicts": ("fail", "pass"), "outcome": "passed", "artifact": "migration plan with rollback verification", } assert result.per_attempt_verdicts == ("fail", "pass") assert result.outcome == "passed" assert result.attempt_log[1].rerun_context == ( "Validation failed: rollback verification is missing. Revise only this failure." ) assert result.artifact == "migration plan with rollback verification" assert result.to_report() == expected_report assert result.rerun_context_fed_back def stubborn_work_fn(rerun_context: str | None) -> str: return "migration plan without rollback verification" exhausted_result = run_with_backpressure(stubborn_work_fn, verify_fn, max_retries=2) assert exhausted_result.outcome == "exhausted" assert exhausted_result.attempts == exhausted_result.max_retries + 1 assert exhausted_result.per_attempt_verdicts == ("fail", "fail", "fail") ``` CrewAI task guardrails are the canonical instance for this card. A task can define guardrails and `guardrail_max_retries`; failed guardrails format validation error context, rerun the agent with that context, and raise when retry exhaustion is reached. The test suite covers both failure-then-success and max-retry exhaustion. AutoGen's writer and critic loop is a partial instance. The critic's `APPROVE` token controls termination, and feedback travels through conversational turns rather than typed rerun context. Dify's provider cooldown is analogy-grade: provider failures cool down a route and push the next attempt elsewhere, but that is infrastructure backpressure rather than agent-output backpressure. ## Determinism Move Backpressure constrains `criteria_drift` by pinning the pass criterion and retry budget in code. A rerun is not driven by the model's sense that the output might be good enough now; it is driven by a recorded verifier result and a counted budget. It also constrains `judge_subjectivity` when the backpressure signal comes from a real verifier: a guardrail, comparator, executable check, or adversarial rubric with explicit failure fields. A subjective "try again" prompt is not backpressure. The determinism move is making the failure a counted, fed-back fact, so the rerun is caused by a recorded check result and a budget, not a retry impulse. ## Observable Signal Every Backpressure report should include: * attempts; * maximum retries; * rerun_context_fed_back boolean; * last failure reason; * per-attempt verdicts; * outcome, such as `passed` or `exhausted`; * artifact or escalation target. A useful report shows the feedback loop: ```text attempts: 2 max_retries: 2 rerun_context_fed_back: true last_failure_reason: rollback verification is missing per_attempt_verdicts: [fail, pass] outcome: passed artifact: migration plan with rollback verification ``` ## Failure Modes * **Swallowed Failure:** downstream validation fails, but the pipeline proceeds. Make failure block progress. * **Blind Retry:** the upstream step is retried without feedback context. Format the verifier failure into rerun context. * **Unbounded Retry:** no retry budget exists. Record attempts and enforce `max_retries`. * **Raw-Error Retry:** a traceback or unstructured string is passed upstream. Convert failure into fields the work step can act on. * **Silent Exhaustion:** retry budget is exhausted and the system quietly stops. Raise or route to Escalation Chain. ## Use When Use this pattern when: * a downstream verifier can fail recoverably; * the upstream step can consume corrective feedback; * bounded automatic revision is cheaper than immediate escalation; * a recorded retry trail is useful for audit or debugging; * an Adversary, Comparator, or guardrail can produce a concrete failure reason. ## Do Not Use When Do not reach for Backpressure when: * the failure is non-recoverable and should escalate immediately; * the upstream step cannot consume feedback; * an Executable Analog passes deterministically without revision; * retry cost is unacceptable; * the verifier cannot state what failed. If the failure cannot be turned into structured rerun context, route to Escalation Chain instead of looping. ## Evidence * **Verification Design Principle 6:** the design doc treats executable verification as the strongest verification move; Backpressure makes a failed check drive bounded revision instead of being swallowed. * **Verification Design Principle 3:** the same design doc frames verification as a step-level practice; Backpressure attaches failure to the step that produced it. * **[CrewAI](https://github.com/crewAIInc/crewAI) task guardrails:** the orchestration sweep records a direct Backpressure instance: failed guardrails format validation error into rerun context, retry within `guardrail_max_retries`, and raise on exhaustion, with tests for both success-after-retry and max-retry failure. * **[AutoGen](https://github.com/microsoft/autogen) writer and critic loop:** the orchestration sweep records a partial instance where critic feedback keeps the writer/critic loop running until explicit approval. * **[Dify](https://github.com/langgenius/dify) provider cooldown:** the orchestration sweep records analogy-grade infrastructure backpressure: provider failures cool down the failing route and move subsequent attempts elsewhere. ## Related Patterns * **Escalation Chain:** receives control when the retry budget is exhausted. * **Adversary:** produces critic findings that can become the backpressure signal. * **Admissibility Gate:** defines the failure-search rubric that gives the loop useful failures. * **Comparator:** supplies the named operator that can drive pass/fail backpressure. * **Causal Tag:** tags each retry attempt so the revision trail is auditable. --- Canonical: https://verificationdesign.com/patterns/orchestration/tool-adapter/ Source: ai-design-patterns/cards/tool-adapter.md # Tool Adapter *(Orchestration Pattern)* ## Name **Tool Adapter** Also known as: Tool Wrapper, Schema Adapter, Function Adapter. ## Intent Normalize model-emitted tool calls at a typed boundary: derive or fetch a schema, validate arguments before invocation, call the tool with typed arguments, and return a typed observation. ## Problem Models do not naturally emit Python function calls. They emit text or loosely shaped JSON that a runtime has to translate into code. The fragile version spreads that translation across call sites: * parse arguments out of a prompt with regexes or string splits; * f-string the values into a function call; * trust that the model used the expected names and types; * reformat the observation inline for the next prompt; * duplicate slightly different schemas in each tool call site. That makes the boundary ambiguous. Is the tool contract the prompt text, the regex, the function signature, or the model's last answer? Malformed arguments can reach side-effectful tools before validation. A hand-copied schema can drift from the actual function. Observations return as raw strings and get reinterpreted differently at the next step. `verification_design.md` Principle 6 names the stronger move: use executable checks instead of interpretation. Tool Adapter applies that principle at the tool boundary. Argument validation is an executable check that runs before the side effect. ## Forces * **Inline conversion vs. typed adapter.** Inline conversion is quick for one script. A typed adapter centralizes the boundary when a model can call tools repeatedly. * **Derived schema vs. hand-written schema.** Deriving from a typed function or fetching a protocol schema reduces drift. Hand-written schemas are sometimes necessary, but they need an owner. * **Validate-before-invoke vs. trust-the-model.** Validation blocks malformed calls before side effects. * **Function adapter vs. protocol adapter.** A local function adapter wraps typed code. A protocol adapter converts external tool schemas into the host framework. * **Strict schema vs. permissive parsing.** Strictness catches bad calls early; permissive parsing can hide boundary errors until the tool runs. ## Solution Put a single adapter at the tool boundary. The adapter owns: * tool registration; * schema derivation or retrieval; * model-facing schema exposure; * argument validation; * invocation with typed arguments; * typed observation formatting; * a call record with schema source, validation result, call id, and return type. The model may propose a tool call. The adapter decides whether the call is valid enough to invoke. ## Mechanism 1. **Register the tool.** Store name, description, callable or protocol reference, and return type. 2. **Derive or fetch the schema.** Use the typed function signature or an external protocol schema. 3. **Expose the schema to the model.** The model sees a stable tool contract rather than prose-only instructions. 4. **Validate arguments before invocation.** Reject missing fields, unexpected fields, or type mismatches before the tool runs. 5. **Return a typed observation.** Record schema source, validated args, call id, observation type, and return value. ## Pattern / Antipattern The same task: allow a model to call `create_ticket(title: str, priority: int)`. The antipattern parses a model string inline and invokes the function without validation. The pattern validates against a derived schema before invocation. ### Antipattern: inline string conversion Ad-hoc conversion is not wrong in a one-off script. It becomes a verification antipattern when an untyped, unvalidated boundary carries model output into side-effectful calls. ```python import re created_tickets: list[dict] = [] def create_ticket(title: str, priority: int) -> dict: created_tickets.append({"title": title, "priority": priority}) return {"ticket_id": "T-1", "title": title, "priority": priority} def inline_tool_call(model_text: str) -> str: title = re.search(r"title=(.*?);", model_text).group(1) priority = re.search(r"priority=(.*)", model_text).group(1) result = create_ticket(title=title, priority=priority) return f"created ticket {result['ticket_id']}: {result['title']}" observation = inline_tool_call("tool=create_ticket title=Prod outage; priority=urgent") ``` The call site has no schema and no validation. The model emitted `priority=urgent`; the tool received it even though the code-facing contract expects an integer. The bug is not regex itself. The bug is that parsing, validation, invocation, and observation formatting are scattered at the call site. ### Pattern: derived schema adapter The structured implementation derives a schema from the function signature and validates model-emitted arguments before invoking the tool. ```python from dataclasses import dataclass from inspect import signature import json from typing import Any, get_type_hints from uuid import uuid4 @dataclass(frozen=True) class CallRecord: call_id: str tool_name: str schema_source: str validated: bool validated_args: dict[str, Any] observation_type: str value: Any validation_error: str | None @dataclass(frozen=True) class Observation: record: CallRecord class FunctionToolAdapter: def __init__(self, name: str, fn, id_factory=None): self.name = name self.fn = fn self.signature = signature(fn) self.type_hints = get_type_hints(fn) self.id_factory = id_factory or (lambda: str(uuid4())) self.calls: list[CallRecord] = [] @property def schema(self) -> dict[str, str]: return { name: getattr(self.type_hints.get(name, Any), "__name__", "Any") for name in self.signature.parameters } def validate(self, args: dict[str, Any]) -> dict[str, Any]: expected = set(self.signature.parameters) if set(args) != expected: raise ValueError(f"expected args {sorted(expected)}, got {sorted(args)}") validated: dict[str, Any] = {} for name, value in args.items(): expected_type = self.type_hints.get(name, Any) if expected_type in {str, int, float, bool}: if type(value) is not expected_type: raise TypeError(f"{name} must be {expected_type.__name__}") elif expected_type is not Any and not isinstance(value, expected_type): raise TypeError(f"{name} must be {expected_type.__name__}") validated[name] = value return validated def run_json(self, args: dict[str, Any]) -> Observation: call_id = self.id_factory() try: validated_args = self.validate(args) except (TypeError, ValueError) as error: self.calls.append( CallRecord( call_id=call_id, tool_name=self.name, schema_source="typed_signature", validated=False, validated_args={}, observation_type="error", value=None, validation_error=str(error), ) ) raise value = self.fn(**validated_args) record = CallRecord( call_id=call_id, tool_name=self.name, schema_source="typed_signature", validated=True, validated_args=validated_args, observation_type=type(value).__name__, value=value, validation_error=None, ) self.calls.append(record) return Observation(record=record) def render_report(adapter: FunctionToolAdapter, record: CallRecord) -> str: validation_error = record.validation_error or "none" rendered_value = ( json.dumps(record.value) if record.value is not None else "none" ) return "\n".join( [ f"tool_name: {record.tool_name}", f"schema_source: {record.schema_source}", f"schema_present: {str(bool(adapter.schema)).lower()}", f"args_validated: {str(record.validated).lower()}", f"validation_error: {validation_error}", f"call_id: {record.call_id}", f"observation_type: {record.observation_type}", f"return_value: {rendered_value}", ] ) created_tickets: list[dict[str, Any]] = [] def create_ticket(title: str, priority: int) -> dict[str, Any]: created_tickets.append({"title": title, "priority": priority}) return {"ticket_id": "T-1", "title": title, "priority": priority} def escalate(ticket_id: str, urgent: bool) -> dict[str, Any]: return {"ticket_id": ticket_id, "urgent": urgent} call_ids = iter(["call-001", "call-002", "call-003"]) tool = FunctionToolAdapter("create_ticket", create_ticket, lambda: next(call_ids)) escalator = FunctionToolAdapter("escalate", escalate, lambda: "escalate-001") observation = tool.run_json({"title": "Prod outage", "priority": 1}) escalation = escalator.run_json({"ticket_id": "T-1", "urgent": True}) try: tool.run_json({"title": "Prod outage", "priority": "urgent"}) except TypeError as error: string_error = str(error) try: tool.run_json({"title": "Prod outage", "priority": True}) except TypeError as error: bool_error = str(error) try: escalator.run_json({"ticket_id": "T-1", "urgent": 1}) except TypeError as error: int_error = str(error) expected_report = """tool_name: create_ticket schema_source: typed_signature schema_present: true args_validated: true validation_error: none call_id: call-001 observation_type: dict return_value: {"ticket_id": "T-1", "title": "Prod outage", "priority": 1}""" expected_rejected_report = """tool_name: create_ticket schema_source: typed_signature schema_present: true args_validated: false validation_error: priority must be int call_id: call-002 observation_type: error return_value: none""" assert tool.schema == {"title": "str", "priority": "int"} assert escalator.schema == {"ticket_id": "str", "urgent": "bool"} assert escalation.record.validated is True assert created_tickets == [{"title": "Prod outage", "priority": 1}] assert string_error == "priority must be int" assert bool_error == "priority must be int" assert int_error == "urgent must be bool" assert [call.validated for call in tool.calls] == [True, False, False] assert tool.calls[1].validation_error == "priority must be int" assert tool.calls[2].validation_error == "priority must be int" assert render_report(tool, observation.record) == expected_report assert render_report(tool, tool.calls[1]) == expected_rejected_report ``` AutoGen's `FunctionTool` and `BaseTool` have this function-adapter shape. `FunctionTool` wraps a Python function, derives an argument model from the typed signature, passes that model to `BaseTool`, exposes a JSON-like schema, validates JSON arguments through the argument model in `run_json`, and only then invokes the wrapped function. CrewAI's MCP resolver is the protocol-adapter sibling. It fetches external MCP tool schemas, converts JSON schemas into local Pydantic argument models where possible, caches schemas by server URL, and wraps the result as a local `BaseTool`. That is the same adapter move applied to a remote tool protocol instead of a Python function signature. The validator above is intentionally small for a pasted example. Production adapters usually delegate this work to an argument model or schema validator. ## Determinism Move Tool Adapter constrains `tool_boundary_ambiguity` by making the model-to-tool boundary a typed, validated schema instead of an implicit string contract. Malformed model output fails at the door instead of inside the tool. It also constrains `criteria_drift` because the schema is a single validation criterion derived from the tool or protocol. The call site does not maintain a hand-copied shape that can drift from the implementation. The determinism move is making the tool contract a derived, validated schema at one boundary. ## Observable Signal Every Tool Adapter report should include: * tool name; * schema source, such as `typed_signature` or `protocol_fetch`; * schema-present boolean; * args-validated boolean; * validation error, or `none`; * call id, deterministic in tests when an id factory is supplied; * observation type; * return value or error surface. A useful report names the boundary: ```text tool_name: create_ticket schema_source: typed_signature schema_present: true args_validated: true validation_error: none call_id: call-001 observation_type: dict return_value: {"ticket_id": "T-1", "title": "Prod outage", "priority": 1} ``` A rejected call should surface the same boundary with a validation failure: ```text tool_name: create_ticket schema_source: typed_signature schema_present: true args_validated: false validation_error: priority must be int call_id: call-002 observation_type: error return_value: none ``` ## Failure Modes * **Inline String Conversion:** arguments are regex-parsed or split out of model text at the call site. Move parsing and validation into an adapter. * **Duplicated Schema:** each call site hand-writes the expected argument shape. Derive the schema from the tool or protocol once. * **Trust-the-Model:** the adapter invokes the tool before validating arguments. Validate before side effects. * **Untyped Observation:** the tool returns a raw string and every caller reformats it differently. Return a typed observation. * **Schema Staleness:** a protocol schema is fetched once and never refreshed. Cache with an invalidation policy or schema version. ## Use When Use this pattern when: * a model calls tools with structured arguments; * tools have typed functions or external protocol schemas; * malformed arguments could cause side effects; * audit requires a single owner for tool format; * observations need to be consumed by later verification or trajectory steps. ## Do Not Use When Do not reach for Tool Adapter when: * there is a single trivial tool with no arguments; * the framework already provides a typed tool boundary and double-wrapping would add no signal; * the wrapper enforces policy rather than adapting types. Use Guardrail Decorator; * the call is fully internal and no model output crosses the boundary. If schema validation is unavailable, label the tool boundary as advisory and add a downstream executable check. ## Evidence * **Verification Design Principle 6:** the design doc treats executable checks as the strongest verification move; schema validation is the executable check at the tool boundary. * **[AutoGen](https://github.com/microsoft/autogen) FunctionTool and BaseTool:** the orchestration sweep records a direct function-adapter instance: typed function signatures generate argument schemas, JSON calls are validated, and validated fields are mapped into the wrapped function. * **[CrewAI](https://github.com/crewAIInc/crewAI) MCP tool resolver:** the orchestration sweep records a direct protocol-adapter instance: external MCP schemas are fetched, converted to local argument models, cached, and exposed as local tools. * **Synthesized inline-conversion antipattern:** no OSS instance was promoted for the antipattern. This card synthesizes the failure case because the issue is the absence of schema ownership and validate-before-invoke behavior. ## Related Patterns * **Guardrail Decorator:** enforces policy at the boundary; Tool Adapter adapts type shape at the same boundary. * **Executable Analog:** schema validation is an executable check for model-emitted tool calls. * **Comparator:** uses an explicit expected-vs-actual check; Tool Adapter applies that check to argument shapes. * **Causal Tag:** tags tool calls so the boundary is auditable across logs. * **Trajectory Cursor:** records tool boundaries as cursor points in the agent trajectory. --- Canonical: https://verificationdesign.com/references/ Source: generated # References - [aaai:36598](https://ojs.aaai.org/index.php/AIES/article/download/36598/38736/40673): Cited from [Verification Design Principles](https://verificationdesign.com/principles/). - [acm:10.1145/2635868.2635920](https://dl.acm.org/doi/10.1145/2635868.2635920): Cited from [Verification Design Principles](https://verificationdesign.com/principles/). - [ADK](https://github.com/google/adk-python): Cited from [Guardrail Decorator](https://verificationdesign.com/patterns/context-and-state/guardrail-decorator/), [Causal Tag](https://verificationdesign.com/patterns/context-and-state/causal-tag/), [Comparator](https://verificationdesign.com/patterns/verification/comparator/), [Judge Harness](https://verificationdesign.com/patterns/verification/judge-harness/), [Admissibility Gate](https://verificationdesign.com/patterns/verification/admissibility-gate/). - [Aider](https://github.com/Aider-AI/aider): Cited from [Delta](https://verificationdesign.com/patterns/verification/delta/). - [arXiv:2212.08073](https://arxiv.org/abs/2212.08073): Cited from [Constitution](https://verificationdesign.com/patterns/context-and-state/constitution/), [Verification Design Principles](https://verificationdesign.com/principles/). - [arXiv:2303.11366](https://arxiv.org/abs/2303.11366): Cited from [Verification Design Principles](https://verificationdesign.com/principles/). - [arXiv:2309.11495](https://arxiv.org/abs/2309.11495): Cited from [Blind Oracle](https://verificationdesign.com/patterns/verification/blind-oracle/), [Verification Design Principles](https://verificationdesign.com/principles/). - [arXiv:2310.01798](https://arxiv.org/abs/2310.01798): Cited from [Constitution](https://verificationdesign.com/patterns/context-and-state/constitution/), [Executable Analog](https://verificationdesign.com/patterns/verification/executable-analog/), [Verification Design Principles](https://verificationdesign.com/principles/), [Home](https://verificationdesign.com/), [About](https://verificationdesign.com/about/). - [arXiv:2410.10934](https://arxiv.org/abs/2410.10934): Cited from [Executable Analog](https://verificationdesign.com/patterns/verification/executable-analog/), [Verification Design Principles](https://verificationdesign.com/principles/). - [arXiv:2504.16828](https://arxiv.org/abs/2504.16828): Cited from [Verification Design Principles](https://verificationdesign.com/principles/). - [arXiv:2511.17826](https://arxiv.org/abs/2511.17826): Cited from [Verification Design Principles](https://verificationdesign.com/principles/). - [arXiv:2512.02304](https://arxiv.org/abs/2512.02304): Cited from [Verification Design Principles](https://verificationdesign.com/principles/). - [arXiv:2512.20845](https://arxiv.org/abs/2512.20845): Cited from [Verification Design Principles](https://verificationdesign.com/principles/). - [arXiv:2601.12294](https://arxiv.org/abs/2601.12294): Cited from [Trajectory Cursor](https://verificationdesign.com/patterns/context-and-state/trajectory-cursor/). - [arXiv:2601.14691](https://arxiv.org/abs/2601.14691): Cited from [Constitution](https://verificationdesign.com/patterns/context-and-state/constitution/), [Judge Harness](https://verificationdesign.com/patterns/verification/judge-harness/). - [arXiv:2603.05399](https://arxiv.org/abs/2603.05399): Cited from [Constitution](https://verificationdesign.com/patterns/context-and-state/constitution/), [Judge Harness](https://verificationdesign.com/patterns/verification/judge-harness/). - [AutoGen](https://github.com/microsoft/autogen): Cited from [Trajectory Cursor](https://verificationdesign.com/patterns/context-and-state/trajectory-cursor/), [Admissibility Gate](https://verificationdesign.com/patterns/verification/admissibility-gate/), [Cross-Family](https://verificationdesign.com/patterns/orchestration/cross-family/), [Adversary](https://verificationdesign.com/patterns/orchestration/adversary/), [Debate](https://verificationdesign.com/patterns/orchestration/debate/), [Escalation Chain](https://verificationdesign.com/patterns/orchestration/escalation-chain/), [Backpressure](https://verificationdesign.com/patterns/orchestration/backpressure/), [Tool Adapter](https://verificationdesign.com/patterns/orchestration/tool-adapter/). - [AutoGPT](https://github.com/Significant-Gravitas/AutoGPT): Cited from [Guardrail Decorator](https://verificationdesign.com/patterns/context-and-state/guardrail-decorator/), [Trajectory Cursor](https://verificationdesign.com/patterns/context-and-state/trajectory-cursor/), [State Baseline](https://verificationdesign.com/patterns/context-and-state/state-baseline/), [Adversary](https://verificationdesign.com/patterns/orchestration/adversary/), [Debate](https://verificationdesign.com/patterns/orchestration/debate/). - [ChatArena](https://github.com/Farama-Foundation/chatarena): Cited from [Cross-Family](https://verificationdesign.com/patterns/orchestration/cross-family/). - [CrewAI](https://github.com/crewAIInc/crewAI): Cited from [State Baseline](https://verificationdesign.com/patterns/context-and-state/state-baseline/), [Escalation Chain](https://verificationdesign.com/patterns/orchestration/escalation-chain/), [Backpressure](https://verificationdesign.com/patterns/orchestration/backpressure/), [Tool Adapter](https://verificationdesign.com/patterns/orchestration/tool-adapter/). - [DeepEval](https://github.com/confident-ai/deepeval): Cited from [Cross-Family](https://verificationdesign.com/patterns/orchestration/cross-family/). - [Dify](https://github.com/langgenius/dify): Cited from [Trajectory Cursor](https://verificationdesign.com/patterns/context-and-state/trajectory-cursor/), [State Baseline](https://verificationdesign.com/patterns/context-and-state/state-baseline/), [Backpressure](https://verificationdesign.com/patterns/orchestration/backpressure/). - [doi:10.1162/tacl_a_00713](https://direct.mit.edu/tacl/article/doi/10.1162/tacl_a_00713/125177): Cited from [Verification Design Principles](https://verificationdesign.com/principles/). - [DSPy](https://github.com/stanfordnlp/dspy): Cited from [Cross-Family](https://verificationdesign.com/patterns/orchestration/cross-family/). - [LangChain](https://github.com/langchain-ai/langchain): Cited from [Causal Tag](https://verificationdesign.com/patterns/context-and-state/causal-tag/), [Blind Oracle](https://verificationdesign.com/patterns/verification/blind-oracle/), [Comparator](https://verificationdesign.com/patterns/verification/comparator/), [Judge Harness](https://verificationdesign.com/patterns/verification/judge-harness/). - [OpenClaw](https://github.com/openclaw/openclaw): Cited from [State Baseline](https://verificationdesign.com/patterns/context-and-state/state-baseline/). - [openreview:4O0v4s3IzY](https://openreview.net/forum?id=4O0v4s3IzY): Cited from [Verification Design Principles](https://verificationdesign.com/principles/). - [openreview:MTvYflAH62](https://openreview.net/forum?id=MTvYflAH62): Cited from [Verification Design Principles](https://verificationdesign.com/principles/). - [the synthesis claim on LLM-judge reliability](https://github.com/verificationdesign/verificationdesign/blob/main/research/synthesis.md#llm-judge-reliability-can-vary-across-benchmarks-and-perturbations-even-when-the-judge-comes-from-a-different-model-family): Cited from [Cross-Family](https://verificationdesign.com/patterns/orchestration/cross-family/).