ENG
ҚАЗ
Paper Review 6: Mixtral of Experts
January 30, 2024
Бұл бет қазақ тілінде әзірленуде.
tags:
LLM
long input transformers
← Paper Review 5: Mistral 7B
Paper Review 7: FlashAttention - Fast and Memory-Efficient Exact Attention with IO-Awareness →