papersSEP 10 04:00 UTC
AgentHijack: Visual Patch Attacks on Multimodal Computer-Use Agents
A new arXiv paper introduces an end-to-end evaluation framework for testing whether a locally placed visual patch can hijack multimodal computer-use agents into executing attacker-chosen commands. Rather than stopping at model-level manipulation, the study checks whether such image-triggered injections lead to verifiable consequences in the agent's operating environment. The work adds to a growing body of research on the security risks of AI agents that control graphical interfaces.