AI Agent Reference Model &
Measurement Framework
6 pillars, 26 components: what to inspect and how to grade it
Know whether your AI agent is ready for the job. 6 pillars. 26 components. One defensible 0–100 score, with a grade for every part you can open and inspect.
The measurement framework: Missing, Measured, Managed
Grade each component 0 (Missing), 1 (Measured), or 2 (Managed). Multiply by weights that sum to 100. The weights are deliberately uneven: governance and mandate are where trust is won or lost, and a long skill list proves little.
Put the framework to work
Get the full AI agent reference model, with all 26 component definitions, artifacts, and measurements, plus the measurement framework scorecard example.
Download the reference model & scorecard
Frequently asked questions
What is the AI agent reference model?
A complete anatomy of a production AI agent: 6 pillars and 26 components covering what an agent is configured with (Agent Mandate, Platform Capabilities, Skills, Tools) and what it executes and produces at runtime (Tasks, Outputs), with Governance spanning all six. Every component has a definition, an artifact you can open, and a measurement, so you can walk the model and check each part.
How does the Missing, Measured, Managed scoring work?
Each of the 26 components earns a grade: 0 (Missing) when nobody can inspect it, 1 (Measured) when the artifact exists and has been checked, even manually, and 2 (Managed) when the measurement is tracked or versioned over time. Grades are multiplied by component weights that sum to 100, producing one comparable 0-100 score per agent.
Why does Governance carry 39 of the 100 points?
Because governance is where trust is won or lost. Its controls are enforced thresholds, and several act as guardrails: they show whether risk stays within tolerance rather than how well the work went. Skills and Tools are both necessary and both easy to verify, which is why they carry the lowest weights in the framework. A comprehensive list of skills and integrations is now table stakes.
How often should an agent be re-scored?
Weekly, on live traffic, against the human baseline: instrument, evaluate, improve, re-test. Drift moves the real score while the surface dashboard stays flat, because model updates, knowledge changes, and shifting traffic decay quality quietly.
Can the framework be adapted by industry or use case?
Yes. The 26 components and their measurements stay consistent across industries; the agents differ. A banking servicing agent, a healthcare scheduling agent, and an IT helpdesk agent each carry their own mandate, policies, tools, and success criteria, and the weights should reflect that work.
Latest agentic AI updates from Druid AI
The AI agent reference model: a complete anatomy of a production agent
Your biggest access channel is still a phone queue. And your automation strategy ignores it.
Scaling Patient Access in Urgent Care: What the Welsh Ambulance Service Learned Deploying Agentic AI
Connect what matters. Make work feel effortless.
See how proven AI agents work for you
Inside real systems, in real scenarios, with accuracy, reliability, and control. So your work feels simpler, not harder.