Accelerating Dropless MoE Training in JAX with NVIDIA Transformer Engine
Mixture of experts (MoE) has become one of the defining architectural trends in large-scale AI model training. DeepSeek, Qwen, and Mixtral are examples of MoE...
关注 Mistral 模型、开放权重与开发者工具。
Mixture of experts (MoE) has become one of the defining architectural trends in large-scale AI model training. DeepSeek, Qwen, and Mixtral are examples of MoE...
模型仓库动态。仓库创建:2025-11-28T17:57:04.000Z。最后修改:2026-09-11T08:37:56.000Z。仓库修改不等同于模型正式发布。 数据经第三方 Hugging Face 镜像采集,需以原始模型页面复核。
模型仓库动态。仓库创建:2026-03-03T10:51:08.000Z。最后修改:2026-09-07T07:28:06.000Z。仓库修改不等同于模型正式发布。 数据经第三方 Hugging Face 镜像采集,需以原始模型页面复核。
模型仓库动态。仓库创建:2026-07-16T10:24:13.000Z。最后修改:2026-08-05T10:06:52.000Z。仓库修改不等同于模型正式发布。 数据经第三方 Hugging Face 镜像采集,需以原始模型页面复核。
模型仓库动态。仓库创建:2026-07-01T16:43:19.000Z。最后修改:2026-07-15T12:24:10.000Z。仓库修改不等同于模型正式发布。 数据经第三方 Hugging Face 镜像采集,需以原始模型页面复核。
模型仓库动态。仓库创建:2026-03-04T16:38:14.000Z。最后修改:2026-07-15T12:23:57.000Z。仓库修改不等同于模型正式发布。 数据经第三方 Hugging Face 镜像采集,需以原始模型页面复核。
模型仓库动态。仓库创建:2026-01-23T13:14:14.000Z。最后修改:2026-07-15T12:23:29.000Z。仓库修改不等同于模型正式发布。 数据经第三方 Hugging Face 镜像采集,需以原始模型页面复核。
模型仓库动态。仓库创建:2026-03-31T09:50:20.000Z。最后修改:2026-07-15T12:23:17.000Z。仓库修改不等同于模型正式发布。 数据经第三方 Hugging Face 镜像采集,需以原始模型页面复核。
模型仓库动态。仓库创建:2025-11-10T07:41:02.000Z。最后修改:2026-07-15T12:22:54.000Z。仓库修改不等同于模型正式发布。 数据经第三方 Hugging Face 镜像采集,需以原始模型页面复核。
模型仓库动态。仓库创建:2025-11-28T18:05:12.000Z。最后修改:2026-07-15T12:22:43.000Z。仓库修改不等同于模型正式发布。 数据经第三方 Hugging Face 镜像采集,需以原始模型页面复核。
模型仓库动态。仓库创建:2025-10-31T08:41:36.000Z。最后修改:2026-07-15T12:22:30.000Z。仓库修改不等同于模型正式发布。 数据经第三方 Hugging Face 镜像采集,需以原始模型页面复核。
模型仓库动态。仓库创建:2025-10-31T08:41:17.000Z。最后修改:2026-07-15T12:22:19.000Z。仓库修改不等同于模型正式发布。 数据经第三方 Hugging Face 镜像采集,需以原始模型页面复核。
模型仓库动态。仓库创建:2025-10-31T08:43:35.000Z。最后修改:2026-07-15T12:22:08.000Z。仓库修改不等同于模型正式发布。 数据经第三方 Hugging Face 镜像采集,需以原始模型页面复核。
模型仓库动态。仓库创建:2025-10-31T08:39:58.000Z。最后修改:2026-07-15T12:21:57.000Z。仓库修改不等同于模型正式发布。 数据经第三方 Hugging Face 镜像采集,需以原始模型页面复核。
模型仓库动态。仓库创建:2025-10-31T08:39:19.000Z。最后修改:2026-07-15T12:21:45.000Z。仓库修改不等同于模型正式发布。 数据经第三方 Hugging Face 镜像采集,需以原始模型页面复核。
模型仓库动态。仓库创建:2025-10-31T08:43:46.000Z。最后修改:2026-07-15T12:21:35.000Z。仓库修改不等同于模型正式发布。 数据经第三方 Hugging Face 镜像采集,需以原始模型页面复核。
# Patch release v5.12.1 Updated the lower bound for PEFT and a fix for auto tokenizer to properly resolve the mistral tokenizer (when `mistral-common` is installed). This is similar to v.5.10.3 minus the fixes that were already included in the main release - v
# vLLM v0.23.0 Release Notes Please note that Minimax M3 is not yet supported in this version. Please follow [vLLM recipe](https://recipes.vllm.ai/MiniMaxAI/MiniMax-M3) for usage guides for M3. ## Highlights This release features 408 commits from 200 contribut
### Summary of major features and improvements * #### More GenAI coverage and framework integrations to minimize code changes * New models supported on CPUs & GPUs: Qwen3 VL * New models supported on CPUs: GPT-OSS 120B * Preview: Introducing the OpenVINO backe
> **_NOTE:_** Please continue using OpenVINO 2025.4 release unless you require the specific bug fixes addressed in the 2025.4.1 version. # Summary of improvements * Preview: Mixture of Experts (MoE) models optimized for CPUs and GPUs, validated for GPT-OSS 20B
### Summary of major features and improvements * #### More GenAI coverage and framework integrations to minimize code changes * **New models supported:** * On CPUs & GPUs: **Qwen3-Embedding-0.6B, Qwen3-Reranker-0.6B, Mistral-Small-24B-Instruct-2501.** * On NP