Reading Notes
19
NLP
18
KV Cache
8
Reading Notes of Random Attention: Rethinking KV Cache Eviction for Efficient Reasoning
Reading Notes of ForesightKV: Optimizing KV Cache Eviction for Reasoning Models by Learning Long-Term Contribution
Reading Notes of RLKV: Which Heads Matter for Reasoning? RL-Guided KV Cache Compression
Reading Notes of SideQuest: Model-Driven KV Cache Management for Long-Horizon Agentic Reasoning
Reading Notes of CacheCraft: Discovering KV Cache Eviction Policies via LLM-Guided Program Evolution
Reading Notes of The Pitfalls of KV Cache Compression
Reading Notes of KVP: Learning to Evict from Key-Value Cache
Reading Notes of H2O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models
LLM Agents
4
Reading Notes of SWE-Pruner: Self-Adaptive Context Pruning for Coding Agents
Reading Notes of EvoDS: Self-Evolving Autonomous Data Science Agent with Skill Learning and Context Management
Reading Notes of ACM: Agentic Context Management for Long Horizon Tasks
Reading Notes of ACON: Optimizing Context Compression for Long-horizon LLM Agents
Efficient Inference
3
Reading Notes of Selective Context: Compressing Context to Enhance Inference Efficiency of Large Language Models
Reading Notes of LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models
Reading Notes of DiffAdapt: Difficulty-Adaptive Reasoning for Token-Efficient LLM Inference