FastE: Readout-Triggered Token Compression for LLM Embedding Inference
A new arXiv paper studies depth-dependent redundancy in the prefix states of final-readout LLM embedding models, including backbones such as Qwen3-Embedding and Qwen3-VL-Embedding. The authors report that dropping these prefix states yields substantial savings, and propose a readout-triggered token compression method named FastE for embedding inference.