Event · 2023-12-11
Open models
Mixtral 8×7B
An open-weight sparse mixture of experts in which only part of the parameters runs for each token.
Mixtral contains several expert blocks, while a router activates only two per layer and token. This increases total capacity without proportional inference cost; open weights and the Apache 2.0 licence allowed researchers and companies to deploy and adapt the system themselves.
Sources
Mistral AI · техническое описаниеOpen primary source Mistral AI · карточка моделиOpen primary source