Back to News
RSS feedarxiv.org

Tool-Using Multimodal LLM Agents Show Weaker Safety Refusal

Summary

A new study examines a safety weakness in agentic multimodal large language models that use tools such as zooming and tagging. Across three safety benchmarks, every top open- and closed-weight MLLM tested was less safe when operating with tools than when answering without them. The relative rate at which models failed to refuse harmful requests increased by as much as 68.7% in the tool-using setting. The researchers base their analysis on more than 100,000 responses, including extended experiments, and propose two possible explanations for the degradation. The findings indicate that adding tool-use capabilities can weaken refusal behavior even in models that otherwise perform strongly on safety evaluations.