HashAttention: Semantic Sparsity for Faster Inference

Leveraging long contexts is crucial for advanced AI systems, but attention computation poses a scalability challenge. While scaled dot-product attention (SDPA) exhibits token sparsity, i.e. only a few pivotal tokens significantly contribute to output, exploiting this sparsity remains challenging. Existing methods either suffer from quality degradation or require substantial additional resources. We show that identifying pivotal tokens is a Maximum Inner Product Search (MIPS) problem. However, existing MIPS solutions are not well-suited for SDPA, as they are not GPU-friendly and often underperform due to the separated query and key distributions. This paper introduces HashAttention, framing pivotal token identification as a recommendation problem. Given a query, HashAttention encodes keys and queries in Hamming space, capturing the required semantic similarity, using learned mapping functions. HashAttention efficiently identifies pivotal tokens for a given query using bitwise operations and computes attention using only these tokens, improving the overall attention efficiency. Trained on generic data, HashAttention reduces tokens used by up to $16\times$ with minimal quality loss, requiring only 32 bits of auxiliary memory per token. Sparsity can be further improved to $32\times$ through task-specific fine-tuning. On A100 GPU, at $32\times$ sparsity, incorporating HashAttention reduces attention latency by up to $4.3\times$ in GPT-FAST and $2.54\times$ in FlashDecode, and achieves up to $3.12\times$ higher throughput for GPT-FAST.

Discussion(0)

No comments yet. Be the first to comment.

Open reviews(0)

Public, signed peer feedback on this preprint.

No reviews yet.

Publication Info

DOI: 10.48550/arxiv.2412.14468
Year: 2024
Published: —
Language: en

Preprint Details

Link Of The Paper: http://arxiv.org/abs/2412.14468

Timeline

Created:June 19, 2026

Related publications

Preprint2024

HashAttention: Semantic Sparsity for Faster Inference

Abstract

Discussion(0)

Open reviews(0)

Related publications

Post-Training Sparse Attention with Double Sparsity

PrefillOnly: An Inference Engine for Prefill-only Workloads in Large Language Model Applications

VSA: Faster Video Diffusion with Trainable Sparse Attention

Accelerating Large-Scale Reasoning Model Inference with Sparse Self-Speculative Decoding

SonicMoE: Accelerating MoE with IO and Tile-aware Optimizations