Home The signal Anatomy of an agent Reference architecture Risk-Tiers Guardrail stack Governance in action Best practices Standards and crosswalk Implementation From the field Roadmap Companion toolkit Straight answers Glossary References
12 · Companion toolkit

Seven artifacts you can use right now

The templates behind everything above, as files rather than screenshots. Open in the browser, or save and adapt. Nothing in them is locked to a vendor, and every threshold is deliberately left blank for you to set.

01 · MD Agent card 02 · CSV Risk scoring 03 · CSV Crosswalk 05 · JSON Event schema SEVEN TEMPLATES · EVERY THRESHOLD LEFT BLANK ON PURPOSE also: threat model, evaluation scorecard, readiness checklist
MD
Agent registration / agent card

The one-page record the registry holds and a reviewer reads first. Identity, accountability, purpose, tier, tools, memory, models, peers, controls, evaluation, operations and retirement plan.

Open artifact 01 →
CSV
Autonomy and risk assessment

Score the use case and each action class on seven dimensions, derive the A1–A4 class and the autonomy ceiling, then read off the required controls. Includes promotion exit criteria you fill in yourself.

Open artifact 02 →
CSV
Standards register and crosswalk

All fifteen AS controls with enforcement mechanism, evidence, and explicit mappings to ISO/IEC 42001, the NIST AI RMF, the AI 600-1 GenAI Profile, EU AI Act articles with provider or deployer role, IMDA dimensions and OWASP ASI risks. Statement-of-applicability rows included.

Open artifact 03 →
MD
Threat-model checklist

Structured on OWASP ASI01–ASI10. Every control is marked hard (deterministic code and configuration) or soft (probabilistic screening), so it is obvious which ones actually hold when the model is wrong.

Open artifact 04 →
JSON SCHEMA
Policy-decision event schema

Draft 2020-12 schema for the record a decision point writes before every action. Carries trace and span ids so it joins your OpenTelemetry data, plus the fields the GenAI conventions do not yet cover: decisions, delegation chains, deterministic limits, approvals, outcomes and retention. Fully worked example included.

Open artifact 05 →
CSV
Evaluation scorecard

Component, step, trajectory and outcome scored separately, plus safety, fairness, cost and judge calibration. Severity-weighted, with a coverage section that asks what is missing rather than only what passed.

Open artifact 06 →
MD
Production readiness checklist

Ten sections, scaled by action risk class A1–A4, with a go/no-go sign-off block. Any unchecked required item blocks release or becomes a time-limited exception with a named owner.

Open artifact 07 →
REPO
All seven, on GitHub

Browse or clone the whole companion folder. Fork it and make the thresholds yours. Corrections and pull requests are welcome.

Open the repository →
How to use these. Start with artifact 02 on one real use case. The scoring conversation surfaces more disagreement about risk appetite than any workshop, and it takes about twenty minutes. Once a use case has a class, the rest of the toolkit tells you which controls it needs and what evidence to keep.

Where would you start?

If you are standing up an agent platform, tightening the controls on one you already have, or preparing for an audit that now includes agents, I am happy to look at it with you.