Defense-as-Skill: Evolving Runtime Guard Skill for Skill-Augmented Agents
Not assessedPlan blockedMeasurementclaim-table8-judge-audit
Human audit of the LLM judges confirms high reliability: the data construction judge achieves 94% balanced audit agreement (88% success precision, 100% failure confirmation) and the final evaluation judge achieves 98% balanced agreement (96% success precision, 100% failure confirmation).
Source: source-paper:Table 8, page 19
Reported and observed measurements
success_precision
m-t8-construction-prec
Reported 88 percentage_points
Observed — percentage_points
failure_confirmation
m-t8-construction-conf
Reported 100 percentage_points
Observed — percentage_points
balanced_audit_agreement
m-t8-construction-agr
Reported 94 percentage_points
Observed — percentage_points
success_precision
m-t8-evaluation-prec
Reported 96 percentage_points
Observed — percentage_points
failure_confirmation
m-t8-evaluation-conf
Reported 100 percentage_points
Observed — percentage_points
balanced_audit_agreement
m-t8-evaluation-agr
Reported 98 percentage_points
Observed — percentage_points
Assessments (0)
No immutable Assessment has been published for this Claim yet.
Runs (0)
No execution has been linked to this Claim yet.