Optimum-Intel v2.2.0 & OpenVINO GenAI 2026.4.0: What's New
Blog post from Hugging Face
Optimum Intel 2.2.0, OpenVINO 2026.4, OpenVINO GenAI 2026.4, and NNCF 3.4 expand Intel’s local AI deployment workflow, which exports Hugging Face models to OpenVINO, quantizes them, and runs optimized inference on Intel CPUs, GPUs, and NPUs. Optimum Intel adds export support for new language, multimodal, OCR, image and video generation, and speech models, including Mistral 3, DeepSeek-OCR-2, Gemma 4, Qwen-Image, LTX-2, and Fun-ASR, while introducing support for speculative decoding methods such as DFlash draft models and Multi-Token Prediction. OpenVINO GenAI adds dedicated image-generation and speech-recognition pipelines, broader vision-language model support including video input, Paged Attention for Mamba 2-based Granite models, faster decoding options, and more detailed performance metrics for vision, audio, and text processing. The release also improves Qwen3-Omni GPU offloading and speech APIs, extends the Node.js API with speech recognition support, and finalizes Mamba 2 OpenVINO representations for Granite-4.0-H models, which can be exported and quantized for local inference.
| Trend | Post Mentions | Total Month Mentions | Posts | Companies | MoM |
|---|---|---|---|---|---|
| Observability | 1 | 472 | 102 | 54 | -85% |
| Vector Search | 1 | 265 | 57 | 33 | -89% |
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.