18 posts in total
2026
Reading Notes of Selective Context: Compressing Context to Enhance Inference Efficiency of Large Language Models
Reading Notes of LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models
Reading Notes of JLens: Verbalizable Representations Form a Global Workspace in Language Models
Reading Notes of Random Attention: Rethinking KV Cache Eviction for Efficient Reasoning
Reading Notes of ForesightKV: Optimizing KV Cache Eviction for Reasoning Models by Learning Long-Term Contribution
Reading Notes of RLKV: Which Heads Matter for Reasoning? RL-Guided KV Cache Compression
Reading Notes of DiffAdapt: Difficulty-Adaptive Reasoning for Token-Efficient LLM Inference
Reading Notes of SWE-Pruner: Self-Adaptive Context Pruning for Coding Agents
Reading Notes of SideQuest: Model-Driven KV Cache Management for Long-Horizon Agentic Reasoning
Reading Notes of CacheCraft: Discovering KV Cache Eviction Policies via LLM-Guided Program Evolution