Does a harness replace model evaluation?
Model evaluation remains part of system validation. A harness also needs tests of authority, access, action execution, human decisions and operating failures.
A practical architecture for agent authority, tool and data scope, pre-action policy, validation, human escalation, monitoring and evidence.
For financial-services platform architects, AI risk leads and security teams.
Last updated: Sep 7, 2026 · Version v1.0 · Not legal advice.
Short answer
A regulated-agent harness connects an agent’s proposed action to explicit authority, bounded access, policy evaluation, validation and human escalation, then records the execution outcome. KLA’s reference architecture organizes that work into seven controls.
Business owner grants authority → agent proposes action → identity and scope checks → pre-action policy → validation or human decision where required → controlled execution → outcome and evidence.
Apply the sequence to each consequential action. Establish network and credential boundaries around it, and send operational failures to the monitoring and remediation process.
KLA’s product mapping below describes relevant surfaces. The acceptance column is a deployment test to perform; it does not assert estate-wide enforcement or completed customer verification.
| Control | KLA mapping | Accountable owner | Acceptance test |
|---|---|---|---|
| Authority | Agent Registry; governed request identity | Platform and business owner | Unrecognized or out-of-scope agent cannot perform the action |
| Tool and data scope | Tool Catalog; Data Boundaries | Platform and data owner | Restricted tool, record and tenant requests are rejected |
| Pre-action policy | Policy Builder; KLA Policy Engine | Policy owner | Exercise allow, warn, require_approval and block on the integrated path |
| Validation | Simulation; workflow-specific checks | Independent reviewer | Invalid output and unsafe state transitions fail before release |
| Human escalation | Decision Desk | Authorized decision owner | Rejection and expiry prevent the approved-action path from executing |
| Monitoring | Assurance Center; Lineage Explorer | Operations and security | A control failure becomes an owned finding with a response |
| Evidence | Audit Trail; Evidence Room | Evidence and records owner | Reconstruct request, decision, human review and actual execution outcome |
List every tool endpoint, credential, data source and delegated agent that can cause the selected action. Route the governed action through the checkpoint and test alternate paths. Record any path that remains outside enforcement.
The host environment owns sandboxing, network isolation and production access. An approval must apply to the concrete action and remain valid at execution; specify behavior when parameters, policy or authority change.
Use a synthetic case with a permitted read, an action requiring approval and a blocked operation. Check downstream state after approval, rejection, expiry and retry. Record identifiers, policy version, reviewer identity and result.
For sealed exports, also run the independent verifier and retain its result. A visible record and a successfully verified bundle provide different evidence; label each accurately.
This is KLA’s architecture proposal, informed by the Bank of England’s harness note. Its seven controls are KLA’s grouping. The accompanying article explains the Bank’s six themes and the note’s scope.
Model evaluation remains part of system validation. A harness also needs tests of authority, access, action execution, human decisions and operating failures.
KLA contributes policy, human decisions and evidence for configured governed paths. Verify integration coverage and separately assign hosting, network, credential and system-level assessment responsibilities.
AI Act standards: agent control mapping
/guides/ai-act-standards-agent-control-mapping
Bank of England harness engineering analysis
/blog/bank-of-england-ai-harness-engineering
Preparing agent controls for EU AI Act standards
/blog/eu-ai-act-standards-agent-controls
SAFR runtime framework
/blog/safr-mas-framework-explained
Discuss your agent control architecture
/book-demo