Skip to content

AI Agent Reference Model &
Measurement Framework

ai_agent_reference_model_and_measurement_framework3

6 pillars, 26 components: what to inspect and how to grade it

Know whether your AI agent is ready for the job. 6 pillars. 26 components. One defensible 0–100 score, with a grade for every part you can open and inspect.

aagent_a_vs_agent_b
GovernanceControls and evidence that keep behavior compliant, auditable, and within risk.
Agent MandateWhy the agent exists and what it may touch: role, targets, handbook, badge, budget.
Platform CapabilitiesMemory, orchestration, and context-rich human handoff.
SkillsPackaged know-how: instructions, templates, scripts.
ToolsControlled interface between probabilistic frontier models and deterministic systems: retrieval (RAG) and tool interfaces.
TasksGoal-bounded work with required context, success criteria, and a terminal state.
OutputsThe response, the system actions taken, and the success rate. Did the work land?
missing_measured_managed

The measurement framework: Missing, Measured, Managed

Grade each component 0 (Missing), 1 (Measured), or 2 (Managed). Multiply by weights that sum to 100. The weights are deliberately uneven: governance and mandate are where trust is won or lost, and a long skill list proves little.

Put the framework to work

Get the full AI agent reference model, with all 26 component definitions, artifacts, and measurements, plus the measurement framework scorecard example.

Frequently asked questions

What is the AI agent reference model?

A complete anatomy of a production AI agent: 6 pillars and 26 components covering what an agent is configured with (Agent Mandate, Platform Capabilities, Skills, Tools) and what it executes and produces at runtime (Tasks, Outputs), with Governance spanning all six. Every component has a definition, an artifact you can open, and a measurement, so you can walk the model and check each part.

How does the Missing, Measured, Managed scoring work?

Each of the 26 components earns a grade: 0 (Missing) when nobody can inspect it, 1 (Measured) when the artifact exists and has been checked, even manually, and 2 (Managed) when the measurement is tracked or versioned over time. Grades are multiplied by component weights that sum to 100, producing one comparable 0-100 score per agent.

Why does Governance carry 39 of the 100 points?

Because governance is where trust is won or lost. Its controls are enforced thresholds, and several act as guardrails: they show whether risk stays within tolerance rather than how well the work went. Skills and Tools are both necessary and both easy to verify, which is why they carry the lowest weights in the framework. A comprehensive list of skills and integrations is now table stakes.

How often should an agent be re-scored?

Weekly, on live traffic, against the human baseline: instrument, evaluate, improve, re-test. Drift moves the real score while the surface dashboard stays flat, because model updates, knowledge changes, and shifting traffic decay quality quietly.

Can the framework be adapted by industry or use case?

Yes. The 26 components and their measurements stay consistent across industries; the agents differ. A banking servicing agent, a healthcare scheduling agent, and an IT helpdesk agent each carry their own mandate, policies, tools, and success criteria, and the weights should reflect that work.

Connect what matters. Make work feel effortless.

See how proven AI agents work for you

Inside real systems, in real scenarios, with accuracy, reliability, and control. So your work feels simpler, not harder.