Understanding KV Cache Compression
A systems view of why KV cache becomes the limiting resource in long-context inference and where compression actually helps.
Read noteKnowledge repository
Deep dives, research surveys, engineering notes, and technology outlooks for building intelligent systems.
A systems view of why KV cache becomes the limiting resource in long-context inference and where compression actually helps.
Read noteA practical model for separating working context, durable knowledge, task state, and execution history.
Read noteA concise map of the compute, memory, and interconnect constraints behind modern language-model serving.
Read note