Defense-as-Skill: Evolving Runtime Guard Skill for Skill-Augmented Agents
尚未评估计划受阻指标结果claim-table8-judge-audit

Human audit of the LLM judges confirms high reliability: the data construction judge achieves 94% balanced audit agreement (88% success precision, 100% failure confirmation) and the final evaluation judge achieves 98% balanced agreement (96% success precision, 100% failure confirmation).

来源:source-paper:Table 8, page 19

报告指标与观测值

success_precision

m-t8-construction-prec

论文报告 88 percentage_points

实际观测 — percentage_points

failure_confirmation

m-t8-construction-conf

论文报告 100 percentage_points

实际观测 — percentage_points

balanced_audit_agreement

m-t8-construction-agr

论文报告 94 percentage_points

实际观测 — percentage_points

success_precision

m-t8-evaluation-prec

论文报告 96 percentage_points

实际观测 — percentage_points

failure_confirmation

m-t8-evaluation-conf

论文报告 100 percentage_points

实际观测 — percentage_points

balanced_audit_agreement

m-t8-evaluation-agr

论文报告 98 percentage_points

实际观测 — percentage_points

Assessment(0)

这条 Claim 暂无已发布的不可变 Assessment。

关联运行(0)

这条 Claim 暂无关联执行。