The Choice Can Be the Attack: Auditing Aligned Backdoors in LLM Agents
Submitted to TACL 2026 with Chowdhury Rakin Haider. The manuscript is not public yet, so this page keeps to the abstract-level summary.
Abstract
LLM agents can be backdoored without visibly failing the task. A trigger can make an agent prefer one valid option over another, such as a brand, vendor, or tool, while the final answer still looks acceptable. SHIFT, the Structured Hidden Influence Test, is a known-trigger audit for this kind of choice steering.
SHIFT reruns matched tasks with and without the trigger, records the valid options and their features, and checks whether the changed choice favors the attacker's target after accounting for ordinary option quality. The current draft positions SHIFT as a practical validation audit for structured choice settings where the auditor can observe the options but not the model internals.