papersSEP 10 04:00 UTC
Paper combines KV cache-aware fine-tuning with recomputation for RAG efficiency
A new arXiv paper tackles the overhead that concatenated retrieved chunks create for KV caches in retrieval-augmented generation systems. The authors fine-tune a model to account for how retrieved passages are joined in the cache while also selectively recomputing cache entries where that still pays off. The work appears under cs.LG with cross-listings in cs.AI and cs.CL.
COVERAGE · 3 REPORTS · LINKS GO TO THE ORIGINAL OUTLETS
arXiv cs.AIFine-Tuning a KV Cache Concatenation-Aware Model or Recomputing KV Caches? Why Not Both? ↗SEP 10 04:00 UTC
arXiv cs.CLFine-Tuning a KV Cache Concatenation-Aware Model or Recomputing KV Caches? Why Not Both? ↗SEP 10 04:00 UTC
arXiv cs.LGFine-Tuning a KV Cache Concatenation-Aware Model or Recomputing KV Caches? Why Not Both? ↗SEP 10 04:00 UTC