← home

Writing

Getting to 10× KV cache compression

Notes on cutting transformer KV cache memory by 10× without sacrificing model accuracy.