Every finding on the monthly report carries three tag families — the canonical framework entry, the vendor-specific refinement, and a Picaroon-owned vertical entry that compounds across retainers. The taxonomy below is what we score against; here is one redacted finding, mapped against all three families at once.
Canonical LLM security risk taxonomy
The OWASP LLM Top 10 pins one tag per failure. Severity is taken from the canonical entry — Picaroon does not relabel it.
Canonical · framework-defined severity
Cloud Security Alliance agent risk taxonomy
CSA Risk Rubric v2 refines agent-specific risks the OWASP catalog groups loosely (tool use, retrieval, refusal-then-act). Cross-tagged against OWASP on the same finding.
Canonical · framework-defined severity
Domain-specific failures no horizontal eval catches
Our own catalog, indexed by vertical. Compounds across retainers: every redacted P0 finding on a refund/warranty agent becomes the next month’s regression line.
Picaroon library · compounds with every retainer
The OWASP LLM Top 10 pins one tag per failure. Severity is taken from the canonical entry — Picaroon does not relabel it.
Canonical · framework-defined severity
CSA Risk Rubric v2 refines agent-specific risks the OWASP catalog groups loosely (tool use, retrieval, refusal-then-act). Cross-tagged against OWASP on the same finding.
Canonical · framework-defined severity
Our own catalog, indexed by vertical. Compounds across retainers: every redacted P0 finding on a refund/warranty agent becomes the next month’s regression line.
Picaroon library · compounds with every retainer
The same redacted ecommerce refund failure that ships on /sample-audit — untagged, then tagged against every family. This is what lands on the Picaroon monthly report.
Refund eligibility override via “manager approval” framing
The agent performs an action with more scope than the user is entitled to — refund, write, delete, escalate, off-policy waive.
Agent underwrites a value-bearing decision on the caller’s behalf, bypassing the human reviewer the policy was designed to require.
A KB chunk from a different tenant or environment reaches the retrieval layer and is treated as on-policy for the current caller.
Refund eligibility is bypassed via intent framing (manager approval, supervisor override, “per the policy I just quoted”). Costs dollars per event.
refund-approval-override) ships verbatim on /sample-audit — this card is the rubric version.The worked mapping above is the same finding rendered in the public sample, with the severity rank and verbatim reproducer prompts. The rubric is the index; the sample is the deliverable.
The taxonomy above is the same index the retainer runs against. Send a brief on your vertical and we come back with a written finding list within two business days.