Defense-as-Skill: Evolving Runtime Guard Skill for Skill-Augmented Agents
Not assessedPlan blockedMeasurementclaim-table8-judge-audit

Human audit of the LLM judges confirms high reliability: the data construction judge achieves 94% balanced audit agreement (88% success precision, 100% failure confirmation) and the final evaluation judge achieves 98% balanced agreement (96% success precision, 100% failure confirmation).

Source: source-paper:Table 8, page 19

Reported and observed measurements

success_precision

m-t8-construction-prec

Reported 88 percentage_points

Observed — percentage_points

failure_confirmation

m-t8-construction-conf

Reported 100 percentage_points

Observed — percentage_points

balanced_audit_agreement

m-t8-construction-agr

Reported 94 percentage_points

Observed — percentage_points

success_precision

m-t8-evaluation-prec

Reported 96 percentage_points

Observed — percentage_points

failure_confirmation

m-t8-evaluation-conf

Reported 100 percentage_points

Observed — percentage_points

balanced_audit_agreement

m-t8-evaluation-agr

Reported 98 percentage_points

Observed — percentage_points

Assessments (0)

No immutable Assessment has been published for this Claim yet.

Runs (0)

No execution has been linked to this Claim yet.