Skip to content

Latest commit

 

History

History
96 lines (67 loc) · 6.48 KB

README.md

File metadata and controls

96 lines (67 loc) · 6.48 KB

Just helping myself keep track of LLM papers that I‘m reading, with an emphasis on inference and model compression.

Transformer Architectures

Foundation Models

Position Encoding

KV Cache

Activation

Pruning

Quantization

Normalization

Sparsity and rank compression

Fine-tuning

Sampling

Scaling

Watermarking

More