Elastic Threshold Attention: Learned Contextual Sparsity for Long-Context Decoding
Introduces Elastic Threshold Attention to optimize KV cache memory usage during long-context decoding.
Massive KV caches can cause severe memory-bandwidth bottlenecks during long-context decoding. Sparse attention methods mitigate this via selective loading, but that comes at a cost: rigid heuristics drop necessary context, leading to quality degradation. We introduce Elastic Threshold Attention (ETA), an end-to-end trainable architecture that achieves hardware-accelerated decoding speed without sacrificing dense mod…