responsible-ai · 2026-10-08 · Tier 2

Tools erode refusals, and agents act before evidence (2026-10-08)

Tools erode refusals, and agents act before evidence (2026-10-08)

Sources: "MLLMs Fail to Refuse when Using Tools Agentically" (NVIDIA / MIT, NeurIPS 2026, arXiv 2610.03938, via @omarsar0, DAIR summary); "From Evidence to Action: How Tool-Using Agents Fail" / SafeActBench (arXiv 2610.07753, raw); agent source-preference study (via @rohanpaul_ai); backdoor persistency in agent post-training (arXiv 2610.07510); AdvSim2Real prompt-injection training (arXiv 2610.08773).

TL;DR. Give a multimodal model tools (zoom, tagging) and it refuses harmful requests less often. Every model tested got worse: up to 68.7% relative rise in refusal failures, 17.7% on average, across MM-SafetyBench, HoliSafe and VLSBench. Claude Opus 4.6, the best plain-chat refuser, went from 13.6% to 18.1% failure. Two causes: tool outputs fill the context and bury the harmful intent ("context dilution"), and the model spends its final turn describing tool results instead of making the safety call ("focus displacement"). Re-inserting the original request and image just before the final answer restores part of the loss. SafeActBench (656 cases) finds the action-side twin: agents judge actions well in static tests but act before required evidence exists in interactive runs.

The safety decision gets buried under tool output
Re-stating the request at the end is a cheap partial fix; plain-chat evals miss the problem.
flowchart LR
  Q["Harmful request<br/><small>image + text</small>"] --> T["Tool calls<br/><small>zoom, tag, search</small>"]
  T --> C["Long context<br/><small>intent diluted</small>"]
  C --> F["Final turn<br/><small>describes tool output</small>"]
  F --> X["Missed refusal<br/><small>up to +68.7%</small>"]
  Q -.->|re-insert| F
  classDef input fill:#d0ebff,stroke:#1971c2,color:#1b1b1b,stroke-width:2px
  classDef loop fill:#fff3bf,stroke:#f08c00,color:#1b1b1b,stroke-width:2px
  classDef core fill:#e5dbff,stroke:#6741d9,color:#1b1b1b,stroke-width:2px
  classDef err fill:#ffe3e3,stroke:#e03131,color:#1b1b1b,stroke-width:2px
  class Q input
  class T loop
  class C,F core
  class X err
  linkStyle 4 stroke:#2f9e44,stroke-width:2px
Blue is the request, amber the tool loop, purple the diluted context, red the failure; the green dashed arrow is the mitigation.

Key points

  • Scope. 100,000+ responses; agent-tuned open models (AdaReasoner), Qwen3.5-122B-A10B, Claude Opus 4.6/4.7, Gemini Agentic Vision.
  • SafeActBench. Ten model-harness setups, six domains, five protocols from static judgment to multi-action workflows. Failures start before execution: agents stop investigating early or act before evidence is established. Once evidence exists, single actions are reliable; multi-step workflows expose unmet prerequisites.
  • Source bias. Agents prefer items by store or site name: with no price listed they assume Walmart is cheaper; 10 of 12 models prefer Booking.com; scholarly search favors arXiv over Medium. Adding the same price to both items cuts the favored store's pick rate by up to 28.3 points.

How this relates to prior wiki pages

  • Same failure shape as KV quantization alignment collapse (09-26), where low-bit KV caches stripped refusals with near-unchanged perplexity. Both show safety behavior is fragile to context changes that capability metrics do not see. Evaluate safety in the deployment configuration.
  • Links to today's harness study (page): enforced harness rules fix stopping where prompts fail; re-inserting the request is the same kind of structural fix for refusals.

Related

Responsible AI · Tool calling · Agent harness engineering