Start with coverage and measured results; open a run only when you need evidence or technical details.
RunMatchRepeat
0/14
claims supported by evidence
0
Supported
0
Challenged or mixed
0
Contradicted
0
Inconclusive
14
Not assessed
Run history
Each row is one recorded execution. Commands, logs, hashes, and signatures are available in its details.
No reproduction runs yet
Runs will appear here as the paper's experiment plans are executed.
Claim–experiment reproduction matrix
See each experiment's execution state, blocker, recovery action, and evidence destination while keeping operations separate from scientific conclusions.
A successful execution does not by itself validate a paper claim
Target state says whether the platform completed the work. The scientific conclusion is determined only by immutable evidence and Assessments. Resource shortages and platform failures are never presented as scientific contradictions.
This task predates target-level records. The states below are conservative projections; no attempts, resource decisions, or evidence links are invented.
SkillSonar generalizes zero-shot to external safety benchmarks, achieving the lowest ASR of 52.90% on SkillSafetyBench and highest safety score of 54.00% on WildClawBench compared to prompt, permission, and allowlist baselines.
Information insufficientclaim-table6-cross-benchmark0 plans0 runs
Scientific conclusionNot assessed
This claim has no executable experiment plan yet.
On benign tasks containing no malicious skill behavior, SkillSonar preserves high utility on Claude Code (0.817 on GLM-5, 0.836 on Claude Haiku 4.5, and 0.724 on GPT-5.4), demonstrating minimal false-positive disruption of safe execution.
Information insufficientclaim-table5-benign-utility0 plans0 runs
Scientific conclusionNot assessed
This claim has no executable experiment plan yet.
Over five matched evaluation rounds, SkillSonar reduces ASR across all individual risk families, achieving reductions of 50.0 pp on Resource & Reliability, 42.9 pp on Execution Safety, 35.3 pp on Specification Integrity, 18.3 pp on Operational Side Effects, 61.5 pp on Privacy & Data Flow (OOD), and 44.8 pp on Capability Control (OOD).
Information insufficientclaim-table10-family-asr0 plans0 runs
Scientific conclusionNot assessed
This claim has no executable experiment plan yet.
Packaging the safety policy as a modular skill-native artifact provides substantial gains over flattening the identical policy into a system prompt: SkillSonar reduces ID ASR from 0.353 to 0.104 and OOD ASR from 0.437 to 0.109, increases ID task utility from 0.655 to 0.815, and lowers token consumption from ~239K to ~188K (a ~21% reduction).
Information insufficientclaim-figure4bc-content-matched0 plans0 runs
Scientific conclusionNot assessed
This claim has no executable experiment plan yet.
Compared to action-level runtime defenses AgentSpec and TS-Guard, SkillSonar achieves superior safety–utility operating points on Claude Code with GLM-5: AgentSpec achieves 0.294 ID ASR but drops task utility to 0.581, while TS-Guard achieves 0.412 ID ASR, whereas SkillSonar reaches 0.104 ID ASR with 0.815 task utility.
Information insufficientclaim-figure5-runtime-guard-baselines0 plans0 runs
Scientific conclusionNot assessed
This claim has no executable experiment plan yet.
Human audit of the LLM judges confirms high reliability: the data construction judge achieves 94% balanced audit agreement (88% success precision, 100% failure confirmation) and the final evaluation judge achieves 98% balanced agreement (96% success precision, 100% failure confirmation).
Information insufficientclaim-table8-judge-audit0 plans0 runs
Scientific conclusionNot assessed
This claim has no executable experiment plan yet.
Under an adaptive attacker performing three iterations of targeted adversarial skill refinement against the final guard, SkillSonar reduces overall ASR from 45.1% to 22.9% (-22.2 percentage points), with reductions across all six individual risk families.
Information insufficientclaim-table9-adaptive-attacker0 plans0 runs
Scientific conclusionNot assessed
This claim has no executable experiment plan yet.
Optimized SkillSonar guards exhibit bidirectional zero-shot transfer across distinct model families without target-model fine-tuning: the GLM-5-optimized guard reduces Kimi-K2.6 ASR to 0.199 (ID) and 0.161 (OOD), while the Kimi-K2.6-optimized guard achieves 0.112 ID ASR on GLM-5, 0.059 on Claude Haiku 4.5, and 0.019 on GPT-5.4.
Information insufficientclaim-figure6-cross-model-transfer0 plans0 runs
Scientific conclusionNot assessed
This claim has no executable experiment plan yet.
In the OpenClaw agent environment, SkillSonar operates without modifying host runtime mechanics, reducing GLM-5 ASR to 0.245 (ID) and 0.232 (OOD), and Claude Haiku 4.5 ASR to 0.265 (ID) and 0.259 (OOD), demonstrating cross-harness portability.
Information insufficientclaim-table4-openclaw-runtime0 plans0 runs
Scientific conclusionNot assessed
This claim has no executable experiment plan yet.
Explicit safety responsibility assignment is necessary for effective runtime guarding: without explicit invocation, ordinary skill selection yields high ASR (0.400 ID, 0.582 OOD), whereas explicit invocation reduces ASR to 0.104 ID and 0.109 OOD while maintaining comparable benign utility (0.756 vs. 0.767).
Information insufficientclaim-figure4a-invocation-ablation0 plans0 runs
Scientific conclusionNot assessed
This claim has no executable experiment plan yet.
During Monte Carlo Tree Search guard evolution, full-evaluation malicious ASR decreases from 0.300 at iteration 1 to 0.078 by iteration 9, demonstrating that feedback-driven refinement effectively optimizes guard policy efficacy.
Information insufficientclaim-figure3-mcts-trajectory0 plans0 runs
Scientific conclusionNot assessed
This claim has no executable experiment plan yet.
On Claude Haiku 4.5 and GPT-5.4, SkillSonar generalizes zero-shot across model architectures, reducing ID/OOD ASR to 0.096/0.254 on Claude Haiku 4.5 and 0.019/0.034 on GPT-5.4, while preserving higher utility than default AcceptEdits.
Information insufficientclaim-table3-cross-model-runtime0 plans0 runs
Scientific conclusionNot assessed
This claim has no executable experiment plan yet.
On repeated runtime evaluation over N=10 runs with GLM-5 on Claude Code, SkillSonar reduces ID ASR from 0.482 to 0.104 and OOD ASR from 0.606 to 0.115 while maintaining 0.779 ID and 0.715 OOD task utility, significantly outperforming system-prompt and permission-based baselines.
Information insufficientclaim-table2-glm5-runtime0 plans0 runs
Scientific conclusionNot assessed
This claim has no executable experiment plan yet.
SCOPE-R contains 206 attack-confirmed malicious skill instances partitioned into 95 train instances (46.1%), 52 ID-test instances (25.2%), and 59 OOD-test instances (28.6%), with Capability Control and Privacy & Data Flow entirely held out from training.
Information insufficientclaim-table1-scoper-splits0 plans0 runs