SkillSonar generalizes zero-shot to external safety benchmarks, achieving the lowest ASR of 52.90% on SkillSafetyBench and highest safety score of 54.00% on WildClawBench compared to prompt, permission, and allowlist baselines.
这条结论尚无可执行实验方案。
优先展示覆盖情况与实测结果;需要审计时再打开运行记录和技术细节。
0/14
条结论获得证据支持
0
获得支持
0
受到挑战或冲突
0
遭到反驳
0
无法判定
14
尚未评估
每一行是一条真实执行记录;命令、日志、哈希和签名都收纳在详情中。
尚无复现运行
论文中的实验方案开始执行后,运行记录会显示在这里。
逐个实验展示当前执行状态、阻塞原因、恢复动作和证据去向;技术执行与科研结论始终分开。
执行成功不等于论文结论成立
目标状态回答平台有没有跑完;右侧科学结论只由不可变证据和 Assessment 决定。资源不足或平台故障不会被写成反驳论文的科研结论。
当前任务创建于目标级记录上线之前。下方状态来自历史任务的保守投影,不会伪造运行尝试、资源决策或证据关系。
SkillSonar generalizes zero-shot to external safety benchmarks, achieving the lowest ASR of 52.90% on SkillSafetyBench and highest safety score of 54.00% on WildClawBench compared to prompt, permission, and allowlist baselines.
这条结论尚无可执行实验方案。
On benign tasks containing no malicious skill behavior, SkillSonar preserves high utility on Claude Code (0.817 on GLM-5, 0.836 on Claude Haiku 4.5, and 0.724 on GPT-5.4), demonstrating minimal false-positive disruption of safe execution.
这条结论尚无可执行实验方案。
Over five matched evaluation rounds, SkillSonar reduces ASR across all individual risk families, achieving reductions of 50.0 pp on Resource & Reliability, 42.9 pp on Execution Safety, 35.3 pp on Specification Integrity, 18.3 pp on Operational Side Effects, 61.5 pp on Privacy & Data Flow (OOD), and 44.8 pp on Capability Control (OOD).
这条结论尚无可执行实验方案。
Packaging the safety policy as a modular skill-native artifact provides substantial gains over flattening the identical policy into a system prompt: SkillSonar reduces ID ASR from 0.353 to 0.104 and OOD ASR from 0.437 to 0.109, increases ID task utility from 0.655 to 0.815, and lowers token consumption from ~239K to ~188K (a ~21% reduction).
这条结论尚无可执行实验方案。
Compared to action-level runtime defenses AgentSpec and TS-Guard, SkillSonar achieves superior safety–utility operating points on Claude Code with GLM-5: AgentSpec achieves 0.294 ID ASR but drops task utility to 0.581, while TS-Guard achieves 0.412 ID ASR, whereas SkillSonar reaches 0.104 ID ASR with 0.815 task utility.
这条结论尚无可执行实验方案。
Human audit of the LLM judges confirms high reliability: the data construction judge achieves 94% balanced audit agreement (88% success precision, 100% failure confirmation) and the final evaluation judge achieves 98% balanced agreement (96% success precision, 100% failure confirmation).
这条结论尚无可执行实验方案。
Under an adaptive attacker performing three iterations of targeted adversarial skill refinement against the final guard, SkillSonar reduces overall ASR from 45.1% to 22.9% (-22.2 percentage points), with reductions across all six individual risk families.
这条结论尚无可执行实验方案。
Optimized SkillSonar guards exhibit bidirectional zero-shot transfer across distinct model families without target-model fine-tuning: the GLM-5-optimized guard reduces Kimi-K2.6 ASR to 0.199 (ID) and 0.161 (OOD), while the Kimi-K2.6-optimized guard achieves 0.112 ID ASR on GLM-5, 0.059 on Claude Haiku 4.5, and 0.019 on GPT-5.4.
这条结论尚无可执行实验方案。
In the OpenClaw agent environment, SkillSonar operates without modifying host runtime mechanics, reducing GLM-5 ASR to 0.245 (ID) and 0.232 (OOD), and Claude Haiku 4.5 ASR to 0.265 (ID) and 0.259 (OOD), demonstrating cross-harness portability.
这条结论尚无可执行实验方案。
Explicit safety responsibility assignment is necessary for effective runtime guarding: without explicit invocation, ordinary skill selection yields high ASR (0.400 ID, 0.582 OOD), whereas explicit invocation reduces ASR to 0.104 ID and 0.109 OOD while maintaining comparable benign utility (0.756 vs. 0.767).
这条结论尚无可执行实验方案。
During Monte Carlo Tree Search guard evolution, full-evaluation malicious ASR decreases from 0.300 at iteration 1 to 0.078 by iteration 9, demonstrating that feedback-driven refinement effectively optimizes guard policy efficacy.
这条结论尚无可执行实验方案。
On Claude Haiku 4.5 and GPT-5.4, SkillSonar generalizes zero-shot across model architectures, reducing ID/OOD ASR to 0.096/0.254 on Claude Haiku 4.5 and 0.019/0.034 on GPT-5.4, while preserving higher utility than default AcceptEdits.
这条结论尚无可执行实验方案。
On repeated runtime evaluation over N=10 runs with GLM-5 on Claude Code, SkillSonar reduces ID ASR from 0.482 to 0.104 and OOD ASR from 0.606 to 0.115 while maintaining 0.779 ID and 0.715 OOD task utility, significantly outperforming system-prompt and permission-based baselines.
这条结论尚无可执行实验方案。
SCOPE-R contains 206 attack-confirmed malicious skill instances partitioned into 95 train instances (46.1%), 52 ID-test instances (25.2%), and 59 OOD-test instances (28.6%), with Capability Control and Privacy & Data Flow entirely held out from training.
这条结论尚无可执行实验方案。