Defense-as-Skill: Evolving Runtime Guard Skill for Skill-Augmented Agents
Not assessedPlan blockedFindingclaim-figure4bc-content-matched

Packaging the safety policy as a modular skill-native artifact provides substantial gains over flattening the identical policy into a system prompt: SkillSonar reduces ID ASR from 0.353 to 0.104 and OOD ASR from 0.437 to 0.109, increases ID task utility from 0.655 to 0.815, and lowers token consumption from ~239K to ~188K (a ~21% reduction).

Source: source-paper:Figure 4(b-c) and Section 5.3 prose, page 11

Reported and observed measurements

attack_success_rate

m-fig4bc-flat-asr-id

Reported 0.353 fraction

Observed — fraction

attack_success_rate

m-fig4bc-flat-asr-ood

Reported 0.437 fraction

Observed — fraction

attack_success_rate

m-fig4bc-skill-asr-id

Reported 0.104 fraction

Observed — fraction

attack_success_rate

m-fig4bc-skill-asr-ood

Reported 0.109 fraction

Observed — fraction

task_utility

m-fig4bc-flat-taskutil-id

Reported 0.655 score

Observed — score

task_utility

m-fig4bc-flat-taskutil-ood

Reported 0.672 score

Observed — score

task_utility

m-fig4bc-skill-taskutil-id

Reported 0.815 score

Observed — score

task_utility

m-fig4bc-skill-taskutil-ood

Reported 0.789 score

Observed — score

benign_task_utility

m-fig4bc-flat-benign-util

Reported 0.748 score

Observed — score

benign_task_utility

m-fig4bc-skill-benign-util

Reported 0.763 score

Observed — score

token_usage

m-fig4bc-flat-tokens

Reported 239 thousands_of_tokens

Observed — thousands_of_tokens

token_usage

m-fig4bc-skill-tokens

Reported 188 thousands_of_tokens

Observed — thousands_of_tokens

Assessments (0)

No immutable Assessment has been published for this Claim yet.

Runs (0)

No execution has been linked to this Claim yet.