
ישראל מירסקי
Token Mines
A Defense Against Agents and Large Language Models
Large Language Models (LLMs) are increasingly being used to automate deepfakes calls and perform social engineering attacks. While current research focuses on designing robust safeguards, there is a critical gap in developing defenses capable of generalizing to a variety of models -and to models that have no safeguards from the start (e.g., finetuned source models). In this work, we explore "token mines"- a sequence of tokens that, if consumed by an LLM, corrupt its hidden state and cause the model to collapse. These token mines can be deployed in a conversation to corrupt a model and thus prove to the victim that he or she is talking with an AI. To create these mines, we optimize a model's output to maximize entropy, resulting in a sequence that not only collapses the base model but also other models as well (i.e., transfers). We evaluate these mines across seven open models spanning GPT-2, LLaMA, BLOOM, Qwen, and Phi architectures with vocabularies ranging from 32K to 250K tokens. Our optimized token sequences achieve 89-96% normalized self-optimization entropy and induce degenerate outputs - including repetition loops, garbage token floods, and cross-lingual artifacts - with over 95% corruption rate in cross-model transfer tests. These findings demonstrate that long-tail vulnerabilities are broadly shared across architectures, suggesting that Token Mines can serve as a viable "universal barrier"for proactively neutralizing unauthorized agents.
| שפת פרסום | אנגלית |
| דפים | 24-30 |
| סטטוס פרסום | פורסם - 31.05.2026 |