Defense-as-Skill: Evolving Runtime Guard Skill for Skill-Augmented Agents

公开
作者:Xiaofang YangZiqi MiaoDianbo SuiJing ShaoLijun Li
更多
复制仓库链接

复现

优先展示覆盖情况与实测结果;需要审计时再打开运行记录和技术细节。

RunMatchRepeat

0/14

条结论获得证据支持

0

获得支持

0

受到挑战或冲突

0

遭到反驳

0

无法判定

14

尚未评估

运行记录

每一行是一条真实执行记录;命令、日志、哈希和签名都收纳在详情中。

尚无复现运行

论文中的实验方案开始执行后,运行记录会显示在这里。

结论—实验复现矩阵

逐个实验展示当前执行状态、阻塞原因、恢复动作和证据去向;技术执行与科研结论始终分开。

当前任务创建于目标级记录上线之前。下方状态来自历史任务的保守投影,不会伪造运行尝试、资源决策或证据关系。

SkillSonar generalizes zero-shot to external safety benchmarks, achieving the lowest ASR of 52.90% on SkillSafetyBench and highest safety score of 54.00% on WildClawBench compared to prompt, permission, and allowlist baselines.

信息不足claim-table6-cross-benchmark0 个方案0 次运行
科学结论尚未评估

这条结论尚无可执行实验方案。

On benign tasks containing no malicious skill behavior, SkillSonar preserves high utility on Claude Code (0.817 on GLM-5, 0.836 on Claude Haiku 4.5, and 0.724 on GPT-5.4), demonstrating minimal false-positive disruption of safe execution.

信息不足claim-table5-benign-utility0 个方案0 次运行
科学结论尚未评估

这条结论尚无可执行实验方案。

Over five matched evaluation rounds, SkillSonar reduces ASR across all individual risk families, achieving reductions of 50.0 pp on Resource & Reliability, 42.9 pp on Execution Safety, 35.3 pp on Specification Integrity, 18.3 pp on Operational Side Effects, 61.5 pp on Privacy & Data Flow (OOD), and 44.8 pp on Capability Control (OOD).

信息不足claim-table10-family-asr0 个方案0 次运行
科学结论尚未评估

这条结论尚无可执行实验方案。

Packaging the safety policy as a modular skill-native artifact provides substantial gains over flattening the identical policy into a system prompt: SkillSonar reduces ID ASR from 0.353 to 0.104 and OOD ASR from 0.437 to 0.109, increases ID task utility from 0.655 to 0.815, and lowers token consumption from ~239K to ~188K (a ~21% reduction).

信息不足claim-figure4bc-content-matched0 个方案0 次运行
科学结论尚未评估

这条结论尚无可执行实验方案。

Compared to action-level runtime defenses AgentSpec and TS-Guard, SkillSonar achieves superior safety–utility operating points on Claude Code with GLM-5: AgentSpec achieves 0.294 ID ASR but drops task utility to 0.581, while TS-Guard achieves 0.412 ID ASR, whereas SkillSonar reaches 0.104 ID ASR with 0.815 task utility.

信息不足claim-figure5-runtime-guard-baselines0 个方案0 次运行
科学结论尚未评估

这条结论尚无可执行实验方案。

Human audit of the LLM judges confirms high reliability: the data construction judge achieves 94% balanced audit agreement (88% success precision, 100% failure confirmation) and the final evaluation judge achieves 98% balanced agreement (96% success precision, 100% failure confirmation).

信息不足claim-table8-judge-audit0 个方案0 次运行
科学结论尚未评估

这条结论尚无可执行实验方案。

Under an adaptive attacker performing three iterations of targeted adversarial skill refinement against the final guard, SkillSonar reduces overall ASR from 45.1% to 22.9% (-22.2 percentage points), with reductions across all six individual risk families.

信息不足claim-table9-adaptive-attacker0 个方案0 次运行
科学结论尚未评估

这条结论尚无可执行实验方案。

Optimized SkillSonar guards exhibit bidirectional zero-shot transfer across distinct model families without target-model fine-tuning: the GLM-5-optimized guard reduces Kimi-K2.6 ASR to 0.199 (ID) and 0.161 (OOD), while the Kimi-K2.6-optimized guard achieves 0.112 ID ASR on GLM-5, 0.059 on Claude Haiku 4.5, and 0.019 on GPT-5.4.

信息不足claim-figure6-cross-model-transfer0 个方案0 次运行
科学结论尚未评估

这条结论尚无可执行实验方案。

In the OpenClaw agent environment, SkillSonar operates without modifying host runtime mechanics, reducing GLM-5 ASR to 0.245 (ID) and 0.232 (OOD), and Claude Haiku 4.5 ASR to 0.265 (ID) and 0.259 (OOD), demonstrating cross-harness portability.

信息不足claim-table4-openclaw-runtime0 个方案0 次运行
科学结论尚未评估

这条结论尚无可执行实验方案。

Explicit safety responsibility assignment is necessary for effective runtime guarding: without explicit invocation, ordinary skill selection yields high ASR (0.400 ID, 0.582 OOD), whereas explicit invocation reduces ASR to 0.104 ID and 0.109 OOD while maintaining comparable benign utility (0.756 vs. 0.767).

信息不足claim-figure4a-invocation-ablation0 个方案0 次运行
科学结论尚未评估

这条结论尚无可执行实验方案。

During Monte Carlo Tree Search guard evolution, full-evaluation malicious ASR decreases from 0.300 at iteration 1 to 0.078 by iteration 9, demonstrating that feedback-driven refinement effectively optimizes guard policy efficacy.

信息不足claim-figure3-mcts-trajectory0 个方案0 次运行
科学结论尚未评估

这条结论尚无可执行实验方案。

On Claude Haiku 4.5 and GPT-5.4, SkillSonar generalizes zero-shot across model architectures, reducing ID/OOD ASR to 0.096/0.254 on Claude Haiku 4.5 and 0.019/0.034 on GPT-5.4, while preserving higher utility than default AcceptEdits.

信息不足claim-table3-cross-model-runtime0 个方案0 次运行
科学结论尚未评估

这条结论尚无可执行实验方案。

On repeated runtime evaluation over N=10 runs with GLM-5 on Claude Code, SkillSonar reduces ID ASR from 0.482 to 0.104 and OOD ASR from 0.606 to 0.115 while maintaining 0.779 ID and 0.715 OOD task utility, significantly outperforming system-prompt and permission-based baselines.

信息不足claim-table2-glm5-runtime0 个方案0 次运行
科学结论尚未评估

这条结论尚无可执行实验方案。

SCOPE-R contains 206 attack-confirmed malicious skill instances partitioned into 95 train instances (46.1%), 52 ID-test instances (25.2%), and 59 OOD-test instances (28.6%), with Capability Control and Privacy & Data Flow entirely held out from training.

信息不足claim-table1-scoper-splits0 个方案0 次运行
科学结论尚未评估

这条结论尚无可执行实验方案。