## Highlights - **Major performance work for attention-heavy LLMs.** - FlashAttention decode kernels were fused and extended for any sequence length ([#28389](https://github.com/microsoft/onnxruntime/pull/28389)). - FlashAttention prefill shared-memory path wa
# vLLM v0.26.0 Release Notes ## Highlights This release features 411 commits from 212 contributors (61 new)! * **New Inkling model family** with a full support stack: base modeling (#48799), piecewise CUDA graph support (#48822), Hopper FA4 relative attention
# Highlights *574 PRs from 169 contributors.* **DSpark: confidence-driven speculative decoding**: A new speculative algorithm. It drafts semi-autoregressively in blocks, then sizes each verify window from the draft's own confidence instead of a fixed draft len
模型仓库动态。仓库创建:2026-06-26T08:39:00.000Z。最后修改:2026-07-22T11:58:47.000Z。仓库修改不等同于模型正式发布。 数据经第三方 Hugging Face 镜像采集,需以原始模型页面复核。 任务:automatic-speech-recognition。
模型仓库动态。仓库创建:2026-06-26T08:34:21.000Z。最后修改:2026-07-22T11:58:31.000Z。仓库修改不等同于模型正式发布。 数据经第三方 Hugging Face 镜像采集,需以原始模型页面复核。 任务:automatic-speech-recognition。
模型动态智谱 GLM 模型仓库(镜像)社区 / 第三方 模型仓库动态。仓库创建:2025-06-28T14:24:10.000Z。最后修改:2026-07-22T11:38:50.000Z。仓库修改不等同于模型正式发布。 数据经第三方 Hugging Face 镜像采集,需以原始模型页面复核。 任务:image-text-to-text。
## What's Changed * Add AGENTS.md by @comfyanonymous in https://github.com/Comfy-Org/ComfyUI/pull/14696 * Add some more stuff to AGENTS.md by @comfyanonymous in https://github.com/Comfy-Org/ComfyUI/pull/14704 * Fix Qwen3-VL tokenizer crash with custom embeddin
模型动态智谱 GLM 模型仓库(镜像)社区 / 第三方 模型仓库动态。仓库创建:2026-06-09T07:18:27.000Z。最后修改:2026-07-15T12:11:45.000Z。仓库修改不等同于模型正式发布。 数据经第三方 Hugging Face 镜像采集,需以原始模型页面复核。 任务:image-to-video。
# vLLM v0.25.1 ## Highlights This release features 2 commits from 2 contributors (1 new)! v0.25.1 is a patch release containing two targeted bug fixes on top of v0.25.0. ### Bug Fixes * **Avoid blocking model launching when no system FFmpeg is available for To
v0.5.15.post1 includes a few patches, mostly for GLM 5.2 - #30454 #30627: Fix DSA model launching on non Cuda/HIP devices - #30858: Fix flashinfer dependency on Cuda 12 images - #31001: Fix NaN outputs caused by flashinfer trtllm FP4 MoE kernels on long input
# vLLM v0.25.0 Release Notes ## Highlights This release features 558 commits from 232 contributors (64 new)! * **Model Runner V2 is now the default for all dense models** (#44443). Building on quantized-model support from the previous release, MRv2 is now the
# Highlights **GLM-5.2 NVFP4, tuned for production**: We took time this cycle to tune GLM-5.2 NVFP4 on Blackwell for optimized production serving. It now runs at **500+ tok/s/user on 8x B300, 450 on 4x GB300** (bs=1). Run GLM-5.2 with our [cookbook](https://do
前沿研究Microsoft Research官方来源 Short chart specifications are easy to write, but often produce uninspiring results. Flint is an open-source visualization language that offers a middle path, letting AI agents create expressive charts from compact, human-editable specifications. The post Flin
模型动态DeepSeek 模型仓库(镜像)社区 / 第三方 模型仓库动态。仓库创建:2026-06-27T02:27:36.000Z。最后修改:2026-07-04T03:15:12.000Z。仓库修改不等同于模型正式发布。 数据经第三方 Hugging Face 镜像采集,需以原始模型页面复核。 任务:text-generation。
模型动态DeepSeek 模型仓库(镜像)社区 / 第三方 模型仓库动态。仓库创建:2026-06-27T03:02:56.000Z。最后修改:2026-07-04T03:14:46.000Z。仓库修改不等同于模型正式发布。 数据经第三方 Hugging Face 镜像采集,需以原始模型页面复核。 任务:text-generation。
Agent 与开发工具Transformers 发布官方来源 # Release v5.13.0 ## New Model additions ### KimiK 2.5, 2.6, and 2.7 This release includes the architecture for Kimi 2.5 which is used by 2.5-2.7: Kimi K2.5 is an open-source, native multimodal agentic model that advances practical capabilities in long-horizon
# vLLM v0.24.0 Release Notes ## Highlights This release features 571 commits from 256 contributors (77 new)! * **MiniMax-M3**: Added support for the new **MiniMax-M3** model (#45381), with a fast follow-on of BF16/FP8 indexer via MSA (#45892), MXFP4 support (#
Agent 与开发工具AgentScope 发布官方来源 ## Major New Features **RAG Module** - Introduced a native RAG (Retrieval-Augmented Generation) module with full support for distributed, multi-tenant, and multi-session RAG services, significantly enhancing knowledge retrieval capabilities. **Long-Term Memory
模型动态DeepSeek 模型仓库(镜像)社区 / 第三方 模型仓库动态。仓库创建:2026-06-28T12:37:17.000Z。最后修改:2026-06-28T12:38:12.000Z。仓库修改不等同于模型正式发布。 数据经第三方 Hugging Face 镜像采集,需以原始模型页面复核。
模型动态DeepSeek 模型仓库(镜像)社区 / 第三方 模型仓库动态。仓库创建:2026-06-28T12:36:20.000Z。最后修改:2026-06-28T12:37:15.000Z。仓库修改不等同于模型正式发布。 数据经第三方 Hugging Face 镜像采集,需以原始模型页面复核。
模型动态DeepSeek 模型仓库(镜像)社区 / 第三方 模型仓库动态。仓库创建:2026-06-28T12:35:29.000Z。最后修改:2026-06-28T12:36:18.000Z。仓库修改不等同于模型正式发布。 数据经第三方 Hugging Face 镜像采集,需以原始模型页面复核。
模型动态DeepSeek 模型仓库(镜像)社区 / 第三方 模型仓库动态。仓库创建:2026-06-28T12:34:59.000Z。最后修改:2026-06-28T12:35:27.000Z。仓库修改不等同于模型正式发布。 数据经第三方 Hugging Face 镜像采集,需以原始模型页面复核。
模型动态DeepSeek 模型仓库(镜像)社区 / 第三方 模型仓库动态。仓库创建:2026-06-28T12:33:24.000Z。最后修改:2026-06-28T12:34:57.000Z。仓库修改不等同于模型正式发布。 数据经第三方 Hugging Face 镜像采集,需以原始模型页面复核。
模型动态DeepSeek 模型仓库(镜像)社区 / 第三方 模型仓库动态。仓库创建:2026-06-28T12:31:30.000Z。最后修改:2026-06-28T12:33:22.000Z。仓库修改不等同于模型正式发布。 数据经第三方 Hugging Face 镜像采集,需以原始模型页面复核。
模型动态DeepSeek 模型仓库(镜像)社区 / 第三方 模型仓库动态。仓库创建:2026-06-28T12:30:07.000Z。最后修改:2026-06-28T12:31:28.000Z。仓库修改不等同于模型正式发布。 数据经第三方 Hugging Face 镜像采集,需以原始模型页面复核。
模型动态DeepSeek 模型仓库(镜像)社区 / 第三方 模型仓库动态。仓库创建:2026-06-28T12:29:12.000Z。最后修改:2026-06-28T12:30:04.000Z。仓库修改不等同于模型正式发布。 数据经第三方 Hugging Face 镜像采集,需以原始模型页面复核。
模型动态DeepSeek 模型仓库(镜像)社区 / 第三方 模型仓库动态。仓库创建:2026-06-28T12:27:36.000Z。最后修改:2026-06-28T12:29:10.000Z。仓库修改不等同于模型正式发布。 数据经第三方 Hugging Face 镜像采集,需以原始模型页面复核。
# Highlights New Model Support: [GLM-5.2](https://docs.sglang.io/cookbook/autoregressive/GLM/GLM-5.2), [LiquidAI LFM2.5](https://docs.sglang.io/cookbook/autoregressive/LiquidAI/LFM2.5), [Kimi-K2.7-Code](https://docs.sglang.io/cookbook/autoregressive/Moonshotai
模型仓库动态。仓库创建:2026-06-26T08:42:01.000Z。最后修改:2026-06-26T08:42:38.000Z。仓库修改不等同于模型正式发布。 数据经第三方 Hugging Face 镜像采集,需以原始模型页面复核。 任务:token-classification。
模型仓库动态。仓库创建:2026-06-22T14:49:37.000Z。最后修改:2026-06-25T07:24:16.000Z。仓库修改不等同于模型正式发布。 数据经第三方 Hugging Face 镜像采集,需以原始模型页面复核。 任务:text-generation。