Intel OpenVINO 发布: 2025.4.0
来源摘要
### Summary of major features and improvements * #### More GenAI coverage and framework integrations to minimize code changes * **New models supported:** * On CPUs & GPUs: **Qwen3-Embedding-0.6B, Qwen3-Reranker-0.6B, Mistral-Small-24B-Instruct-2501.** * On NPUs: **Gemma-3-4b-it and Qwen2.5-VL-3B-Instruct.** * Preview: Mixture of Experts (MoE) models optimized for CPUs and GPUs, validated for Qwen3-30B-A3B. * **GenAI pipeline integrations: Qwen3-Embedding-0.6B and Qwen3-Reranker-0.6B for enhanced retrieval/ranking, and Qwen2.5VL-7B for video pipeline.** * #### Broader LLM model support and more model compression techniques * **Gold support for Windows ML\*** enables developers to deploy AI models and applications effortlessly across CPUs, GPUs, and NPUs on Intel® Core™ Ultra processor-powered AI PCs. * **The Neural Network Compression Framework (NNCF) ONNX backend now supports INT8 static post-training quantization (PTQ) and INT8/INT4 weight-only compression** to ensure accuracy parity with OpenVINO IR format models. SmoothQuant algorithm support added for INT8 quantization. * **Accelerated multi-token generation for GenAI, leveraging optimized GPU kernels to deliver faster inference, smarter KV-cache reuse, and scalable LLM performance.** * **GPU plugin updates include improved performance with prefix caching for chat history scenarios and enhanced LLM accuracy with dynamic quantization support for INT8.** * #### More portability and performance to run AI at the edge, in the cloud, or locally. * **Announcing support for Intel® Core™ Ultra Processor Series 3.** * **Encrypted blob format support added for secure model deployment with OpenVINO™ GenAI.** Model weights and artifacts are stored and transmitted in an encrypted format, reducing risks of IP theft during deploym
阅读原始来源- 来源
- Intel OpenVINO 发布 · 官方来源
- 来源发布
- 2025/12/01 21:26
- 来源更新
- 2025/12/18 17:08
- 首次采集
- 2026/09/19 13:21
本文为公开信息索引与摘要,详情及后续变化请以原始来源为准。