Future Proof The Authority Stack
European AI Agent Certification Standard
Methodology v 1.0 · Friday, 17 April 2026
Published by Future Proof Intelligence
Document
Certification Tier Specification
Revision
2026.04
Tiers
Five
Validity
12 Months
Scoring Bands

The rubric at a glance

The final score is a single number between zero and one hundred, produced by the methodology. The number alone is not the certification. The certification is the tier the number falls into, together with the evidence file and the registered expiry date.

To see how evidence becomes a number, and why a weighted total of 58 does not reach the Advanced tier, read the worked example, where one constructed deployment is scored dimension by dimension. No real organisation is scored there, and none has been assessed under this methodology to date.

Pre Assessment
In Progress
Certified
Advanced
Elite
0 to 19 20 to 34 35 to 54 55 to 74 75 to 100
75+

Elite

Exemplar tier

What the score range means

The agent is operating at the top of the framework. Every dimension is materially above the operating floor, including the two highest weighted ones. No agent has been assessed at this tier to date, and the tier is described here as the methodology defines it rather than by example.

Requirements to earn Elite

  • Minimum raw score of eight on every dimension, with no exceptions
  • Independent assurance already performed or scheduled within the current year
  • Board level review of the agent in the last six months
  • Documented incident history showing at least one contained real event
  • Autonomy envelope tied to insurance policy wording

What an insurer learns from Elite

This is the tier where an insurer can price against a known, repeatable risk posture. Agent Certified has no arrangement with any insurer or supervisory authority, and no Elite result has been adopted as an underwriting input.

Validity & recertification

Twelve months. Recertification requires a full re assessment, not a desk update, and must include at least one new piece of incident or drill evidence.

55+

Advanced

Higher exposure ready

What the score range means

The agent is materially above the operating floor and suitable for higher exposure deployments, including use inside regulated workflows where audit trail and human oversight are non negotiable.

Requirements to earn Advanced

  • Minimum raw score of six on every dimension
  • No single dimension below five
  • Governance owner named and referenced in board minutes
  • Regression evaluation suite run on every change
  • Autonomy envelope enforced in code, not only in policy

What an insurer learns from Advanced

A credible risk posture. Advanced agents are acceptable for standard commercial cover, with sector specific endorsements added when deployment scope exceeds internal tooling.

Validity & recertification

Twelve months. Recertification may be partial if the agent has not materially changed, with desk review of updated evidence.

35+

Certified

Operating floor

What the score range means

The agent meets the minimum floor for certification. Governance, oversight and technical controls exist, are documented, and have been inspected. Counterparties may rely on the mark for standard commercial use.

Requirements to earn Certified

  • Minimum raw score of four on every dimension
  • Documented guardrails and a verified kill switch
  • Named accountable owner
  • Incident playbook in writing
  • Autonomy policy published internally

What an insurer learns from Certified

The operator has crossed the threshold below which reliance is not defensible. Certified is the minimum bar most European insurers will begin asking for by the second half of 2026.

Validity & recertification

Twelve months. A light touch annual review is required if no material change has occurred. Any change to the autonomy envelope triggers re assessment regardless of calendar date.

20+

In Progress

Active candidate

What the score range means

The operator has initiated formal controls but the evidence is not yet sufficient for a certification call. The agent is recognised as an active candidate with a known remediation path.

Requirements to earn In Progress

  • Minimum raw score of two on every dimension
  • A written remediation plan with target dates
  • Named owner, even if the role is still interim
  • Acknowledgement by the board that the agent is in production

What an insurer learns from In Progress

The operator is engaged with the framework and is progressing toward certification. This is an acceptable posture for pilots and internal tooling, not for client facing or regulated workflows.

Validity & recertification

Six months. The tier is intentionally short lived. Agents that remain at In Progress after twelve months without movement lose the tier.

<20

Pre Assessment

Baseline

What the score range means

Baseline acknowledged. The operator has engaged with the framework but governance, oversight and technical evidence are not yet sufficient to make a certification call in any direction.

Conditions

  • Evidence file opened with the assessment desk
  • No technical requirement beyond the existence of an agent in production or pilot
  • The operator receives a written gap report against the framework

What an insurer learns from Pre Assessment

Pre Assessment is a starting position, not a certification. It indicates that the operator has entered the process. Insurers typically treat Pre Assessment status as equivalent to uncertified for the purpose of reliance.

Validity & recertification

Three months. Pre Assessment lapses automatically unless the operator progresses to In Progress or above.

How a tier is arrived at is set out in the assessment process step by step, and the scoring logic behind the tiers is explained in the complete Agent Certified framework guide. The dimension that most often decides whether an agent clears Advanced is treated separately in performance and reliability. Further reading is indexed under AI agent certification.

Reference

Side by side comparison

The five tiers in one table. Use this as a quick read for risk committees, procurement teams and insurer intake forms.

Level Score Minimum per dimension Signal Validity
Elite 75 to 100 8 / 10 Reference profile. Every dimension well above floor. 12 months, full re assessment
Advanced 55 to 74 6 / 10 Regulated workflows and higher exposure. 12 months, desk possible
Certified 35 to 54 4 / 10 Standard commercial use. Operating floor met. 12 months, annual review
In Progress 20 to 34 2 / 10 Candidate. Pilots and internal tooling only. 6 months
Pre Assessment 0 to 19 n/a Entry point. Treated as uncertified for reliance. 3 months
Common Questions

Questions the tier specification answers

How is a composite agent certified?

A composite agent, one assembled from several sub agents, tools or models rather than a single model behind a single endpoint, is certified as one unit of reliance and receives one score. The unit of assessment is the thing a counterparty relies on, not the parts it happens to be built from. In practice that means three additional evidence requirements. The Distribution Control dimension is scored against the widest authority any component holds, not the average, because an attacker or a mistake reaches the whole through the weakest boundary. The Autonomy Envelope is scored against the longest unsupervised chain the composite can execute end to end, including any step where one sub agent invokes another. AI Integration is scored against whether handoffs between components preserve the identity, authority and audit trail of the original request, or quietly launder them. A composite whose components are individually well governed but whose handoffs are not will score below the sum of its parts, which is the intended behaviour.

What should be in an eval suite before an agent goes live?

An evaluation suite that satisfies the Performance and Reliability dimension covers five things, and a suite missing any of them does not score. A held out task set drawn from real production inputs rather than from the examples used to build the agent, with a stated pass threshold agreed before the run. Adversarial cases, including prompt injection attempts against every external data path the agent reads, since prompt injection is ranked LLM01 in the OWASP Top 10 for Large Language Model Applications. Regression coverage, so that the same suite runs against every model or prompt change and results are comparable over time. Failure mode cases that assert what the agent does when a tool is unavailable, a response is malformed or a request falls outside scope, because unhandled failure is where autonomy causes damage. And a recorded run history with dates, versions and results retained, because a suite that exists but has no history proves only that someone wrote it.

Which certification level does an AI agent actually need?

The level follows the exposure, not the ambition. An agent running internal tooling with a human approving anything that leaves the organisation is served by In Progress or Certified, and paying for a higher tier buys nothing a counterparty will read. An agent acting on customers, moving money or producing output a third party relies on should be at Certified as a floor and Advanced where the workflow is regulated, because Advanced sets a minimum of six out of ten in every dimension and so rules out the pattern where a strong governance score conceals a weak oversight design. Elite exists as the top of the scale rather than as a target: it requires eight out of ten in every one of the seven dimensions, which is a description of an unusually mature operation rather than a procurement requirement. The practical question to ask before commissioning an assessment is which tier the counterparty asking for evidence will actually read.

Next Step

See where your agent lands.

The first output of an assessment is a score and a tier placement, with a written evidence file. An assessment runs four to six weeks from intake to certification record.

The Regulatory Context

Read the law that the methodology operates under.

The Agent Certified framework is calibrated to Article 26 of the EU AI Act, the revised Product Liability Directive, and the supervisory expectations of EIOPA and the AI Office. Agent Liability EU is the operator desk on those instruments.

agentliability.eu