ישראל מירסקי

אקדמי בכיר

Token Mines

A Defense Against Agents and Large Language Models

Anton Pasternak, Tomer Ashkenazy, Itay Arad, Roy Weiss, Yisroel Mirsky

Large Language Models (LLMs) are increasingly being used to automate deepfakes calls and perform social engineering attacks. While current research focuses on designing robust safeguards, there is a critical gap in developing defenses capable of generalizing to a variety of models -and to models that have no safeguards from the start (e.g., finetuned source models). In this work, we explore "token mines"- a sequence of tokens that, if consumed by an LLM, corrupt its hidden state and cause the model to collapse. These token mines can be deployed in a conversation to corrupt a model and thus prove to the victim that he or she is talking with an AI. To create these mines, we optimize a model's output to maximize entropy, resulting in a sequence that not only collapses the base model but also other models as well (i.e., transfers). We evaluate these mines across seven open models spanning GPT-2, LLaMA, BLOOM, Qwen, and Phi architectures with vocabularies ranging from 32K to 250K tokens. Our optimized token sequences achieve 89-96% normalized self-optimization entropy and induce degenerate outputs - including repetition loops, garbage token floods, and cross-lingual artifacts - with over 95% corruption rate in cross-model transfer tests. These findings demonstrate that long-tail vulnerabilities are broadly shared across architectures, suggesting that Token Mines can serve as a viable "universal barrier"for proactively neutralizing unauthorized agents.

שפת פרסום אנגלית
דפים 24-30
סטטוס פרסום פורסם - 31.05.2026

ASJC Scopus subject areas

Computer Networks and Communications
Computer Science Applications
Information Systems
Software
גישה למסמך
10.1145/3803629.3803674
קבצים וקישורים אחרים
Link to publication in Scopus