papersTODAY 04:00 UTC
arXiv paper proposes learned selection of poison sets for LLM backdoor attacks
A new arXiv preprint introduces a method that learns which examples to poison in order to make backdoor attacks on fine-tuned language models more effective. The authors note that prior work usually holds the number of poisoned examples fixed, and their approach instead optimizes the choice of poison set. The paper appears in both cs.AI and cs.LG listings.