ENG
ҚАЗ
Paper Review 7: FlashAttention - Fast and Memory-Efficient Exact Attention with IO-Awareness
February 6, 2024
Бұл бет қазақ тілінде әзірленуде.
tags:
efficiency
long input transformers
← Paper Review 6: Mixtral of Experts
Paper Review 8: SGDR: SGD with warm restarts(+ theory behind the Grad descent algorithm) →