ENG
ҚАЗ
Бұл бет қазақ тілінде әзірленуде.
- Paper Review 8: SGDR: SGD with warm restarts(+ theory behind the Grad descent algorithm)
- Paper Review 7: FlashAttention - Fast and Memory-Efficient Exact Attention with IO-Awareness
- Paper Review 6: Mixtral of Experts
- Paper Review 5: Mistral 7B
- Paper Review 4: Self-attention Does Not Need O(n^2) Memory
- Paper Review 3: Pix2Struct is an image-encoder-text-decoder based on the Vision Transformer (ViT)
- Paper Review 2: MATCHA : Enhancing Visual Language Pretraining with Math Reasoning and Chart Derendering
- Paper Review 1: LLaMA - Open and Efficient Foundation Language Models