Home / Companies / Comfy / Blog / August 2026

August 2026 Summaries

13 posts from Comfy

Filter
Month: Year:
Post Summaries Back to Blog
ComfyUI version 0.34.0 adds native support for Trellis.2 and Pixal3D, two open 3D-generation models that create textured 3D assets from a single image, while removing the custom nodes, compiled CUDA extensions, environment downgrades, and non-commercial NVIDIA dependencies previously required. Trellis.2, a 4-billion-parameter model using O-Voxel latent representations, supports detailed geometry, varied topology, and PBR base color, roughness, and metallic materials, while Pixal3D builds on its backbone to provide pixel-aligned geometry with closer correspondence to the input image. The integration also introduces rebuilt 3D loading, previewing, and saving nodes; native mesh-processing tools for remeshing, decimation, smoothing, hole filling, rendering, and painting; and expanded PBR workflows that add UV unwrapping plus baked normal and ambient-occlusion maps. Although closed-source services such as Hunyuan 3D, Tripo, and Rodin are described as producing higher-quality results, the update positions Trellis.2 and Pixal3D as freely available, commercially usable local alternatives for prototyping, stylized work, iteration, and controlled production pipelines on consumer hardware.
Aug 31, 2026 1,244 words in the original blog post.
Google’s Gemini Omni 1.1 Flash is available in ComfyUI through the Gemini Video Omni Partner Node, providing text-to-video, image-to-video, reference-based generation, video editing, scene extension, generated audio, and output resolutions from 360p to 4K. The multimodal model processes text, images, audio, and video together to maintain consistency, supports conversational edits that aim to preserve unaffected parts of a clip, and can extend scenes while retaining motion, lighting, and character continuity. Users can configure resolution, aspect ratio, task type, and optional image or video inputs, with workflows encouraged to iterate at 720p before rendering final versions in 1080p or 4K. Prompt guidance emphasizes detailed descriptions for generation, explicit instructions for continuous shots, concise commands for edits, time-based language for event placement, and direct specifications for soundtrack, readable on-screen text, and exclusions, while image tags can designate opening frames or reference assets.
Aug 28, 2026 849 words in the original blog post.
Comfy has become the first official reseller of MiniMax H3 commercial-use licenses, enabling businesses to use the open-weight video model locally for commercial products, client work, and other professional applications. H3 supports 2K video with native stereo audio and can run on relatively accessible hardware, while licenses purchased through Comfy provide commercial output rights, permission for fine-tuning and LoRA training, client and downstream project use, and coverage for MiniMax audio and music models. Professional and Enterprise plans are available, with Enterprise customers receiving access to all H3 versions, including undistilled weights and future releases; however, the licenses do not permit operating an inference platform or marketplace that sells H3 access. Comfy also offers Forward Deployed Creatives, specialists who can help teams train custom LoRAs, develop production workflows, and transfer the resulting systems to customers, while Comfy Cloud users already receive commercial rights and do not need a separate local-use license.
Aug 27, 2026 512 words in the original blog post.
Alibaba’s Wan 3.0 video generation model is now available in ComfyUI, offering single-pass video creation lasting up to 30 seconds at 480p, 720p, or 1080p and supporting multiple aspect ratios, optional audio, adaptive framing, and automatic duration selection. The model accepts prompts of up to 20,000 characters and can use as many as 20 reference assets, including images, videos, audio clips, files, or URLs, with prompt-based @ references enabling targeted control over elements such as character identity, camera movement, and voice timbre. It also supports document and webpage parsing, instruction-based editing of existing footage, and video extension up to a combined 30-second limit. A new lower-cost 480p option is intended for faster iteration, while users can access the model by updating ComfyUI to version 0.33.4, using Comfy Cloud, or loading Wan 3.0 nodes and workflow templates.
Aug 25, 2026 445 words in the original blog post.
ComfyUI and MiniMax are hosting a fully virtual community challenge from August 20 to September 1, 2026, inviting participants to create videos of up to 90 seconds using MiniMax H3’s native synchronized audio and video generation within ComfyUI. Entries must include the final H3-generated video with synchronized sound and the associated ComfyUI workflow JSON, while separately adding an audio track after generation is prohibited; participants may use reference audio, external tools, and other models provided the core video and audio are generated through H3 in ComfyUI. H3 can be used locally for free on compatible hardware, including optimized support for consumer GPUs, or through credit-based Comfy Cloud, and Comfy MCP allows users to manage or adapt workflows with natural-language agents. One submission per user is permitted, entries must be lawful, safe for work, and free of unlicensed intellectual property or likenesses, and submitted workflows and outputs will be shared with the community for browsing and remixing. Prizes include an RTX 5090 for Best Overall and RTX 5060 Ti GPUs for Best Technical/Workflow, Best Creative, and an MCP-focused bonus category, with winners announced during a September 2 livestream and judged by ComfyUI, MiniMax, guest judges, and community input.
Aug 19, 2026 1,351 words in the original blog post.
ComfyUI has open-sourced Comfy MCP’s local connection, allowing AI agents such as Claude, Codex, and Cursor to interact with a user’s locally installed ComfyUI instance alongside its existing cloud capabilities. Through one Comfy account, agents can assess local GPU hardware, inspect installed models and custom nodes, manage model downloads and local instances, build or modify workflows through natural-language requests, and choose whether tasks are better suited to local hardware or the cloud. The announcement highlights using the system to configure and run demanding open-weight models such as MiniMax H3, with the agent checking compatibility, locating weights and templates, validating available nodes, and saving results locally. Local MCP is positioned as enabling low-cost large batch jobs, integration with desktop tools such as Blender, Houdini, and DaVinci through a shared file system, and automated discovery and installation of custom node packs, with setup instructions and source code available through Comfy’s documentation and GitHub repository.
Aug 18, 2026 716 words in the original blog post.
An August 2026 ranking of eight AI creative workflow platforms evaluates Comfy, Adobe Firefly, Runway, Midjourney, Canva, Leonardo, Krea, and Freepik/Magnific on workflow depth, model flexibility, output quality, control, scalability, pricing, and ecosystem maturity. Written by the Comfy team and explicitly acknowledging that affiliation, it places Comfy first for its open-source, node-based, model-agnostic workflows, reproducible settings, large community-node ecosystem, and local, cloud, and API deployment options, while noting its technical learning curve, hardware demands, and third-party dependency and security concerns. Adobe Firefly is positioned for Creative Cloud-based teams, Runway for browser-based AI video production, Midjourney for visually strong concept art, Canva for accessible template-led marketing content, Leonardo for custom-trained game and brand assets, Krea for real-time multimodal experimentation, and Freepik/Magnific for stock sourcing and generative enhancement. The ranking emphasizes that no single platform fits every production need, arguing that professional teams often combine tools and should select platforms based on whether they require repeatable, inspectable pipelines or faster, more constrained creative workflows.
Aug 15, 2026 3,622 words in the original blog post.
MiniMax Music 3 is an open-weight music generation model available in ComfyUI that produces complete songs of up to five minutes from lyrics and a music description, delivering 32 kHz, 16-bit stereo audio. Designed to maintain long-range musical coherence, it can preserve melody, rhythm, vocal identity, and evolving arrangements across structures such as verses, choruses, bridges, instrumental sections, and outros. Its hybrid architecture combines an 8B global language model for overall song structure, a 0.6B local model for acoustic detail, and continuous Flow Matching and Flow-VAE synthesis components intended to improve vocal articulation, instrumental texture, and temporal continuity. Users can control outputs through tagged lyric sections and Structured Captions specifying metadata such as genre, BPM, key, emotional progression, vocal characteristics, instruments, and section-level arrangement changes. MiniMax also provides a caption-rewriter skill to expand short prompts into structured descriptions, while the model and workflow can be accessed through ComfyUI version 0.33.0 or later, Comfy Cloud, and the published model weights repository.
Aug 13, 2026 607 words in the original blog post.
LTX-2.5, LTX’s latest open video-generation model, is available in ComfyUI at launch and is designed to run quickly on consumer GPUs while offering downloadable weights and fine-tuning support. The release introduces Diffusion Fidelity Rendering, which assigns more computation to visually complex scenes through structured latent generation, adaptive keyframes, and a pixel-diffusion rendering stage, alongside a new diffusion video decoder intended to improve faces, text, and fast-motion detail. It also supports native multi-shot video generation that maintains continuity across cuts, a custom Gemma 4 12B text encoder for longer and more detailed prompts, a lightweight prompt enhancer, and an experimental feature that automatically predicts clip duration. LTX-2.5 retains native 4K output, synchronized audio and video, and frame rates up to 50fps, and is offered as open-weight development and distilled models as well as hosted Fast and Pro Partner Node variants with differing resolution, duration, and frame-rate limits. Users can access it by updating ComfyUI or using Comfy Cloud, downloading the model weights, and loading supplied text-to-video, image-to-video, or related workflow templates.
Aug 12, 2026 628 words in the original blog post.
Seedance 2.5 is available in ComfyUI through Partner Nodes and Comfy Cloud, offering single-pass generation of continuous 30-second videos rather than stitched short clips. The model can use up to 50 reference assets, including images, videos, and audio files, with inline prompt references for controlling characters, settings, movement, and sound. It supports timeline-based prompts for specifying shot timing, camera choices, transitions, and cuts, as well as natural-language editing of existing footage, such as removing subjects, reconstructing backgrounds, or extending a clip. Seedance 2.5 also produces synchronized spoken dialogue in more than ten languages, including clips that switch between languages. Users can access it by updating ComfyUI or using Comfy Cloud, then loading templates or downloadable reference-to-video and text-to-video workflows.
Aug 08, 2026 425 words in the original blog post.
Wan Animate 2 has received native ComfyUI support, introducing an end-to-end character animation system that feeds driving video directly into a redesigned Diffusion Transformer rather than relying on intermediate motion extraction. The approach is intended to improve motion fidelity, preserve character identity more consistently, and retain details such as facial expressions and finger movements, while text prompts can independently control the output camera viewpoint. An efficient Wan Animate 2 Lite version targets real-time streaming animation, and the models also support video extension and long-sequence generation through context windows. ComfyUI includes a WanAnimate2ToVideo conditioning node for reference characters and driving videos, plus an optional WanAnimate2Cache node that reuses pose-processing results to approximately halve generation time at the cost of additional memory. Users can access the feature by updating ComfyUI or using Comfy Cloud, loading provided templates or workflows, connecting source media, configuring resolution and duration, and obtaining model weights from Comfy-Org/Wan-Animate-2.
Aug 08, 2026 427 words in the original blog post.
FLUX 3 Video from Black Forest Labs is now accessible in ComfyUI through Partner Nodes, offering video generation first while additional image capabilities, fast variants, and open model weights are planned. Built as a unified multimodal architecture jointly trained across video, audio, images, and action prediction, the model can generate up to 20-second clips with optional native dialogue, sound effects, and ambience. It supports text-to-video, image-guided video using frames or keyframes, video continuation, multi-scene outputs, multilingual lip-synced dialogue, integrated typography, and chained generations for longer stories. The release emphasizes broad stylistic flexibility beyond conventional cinematic video, ranging from photorealistic imagery to animation, and users can begin by updating ComfyUI or using Comfy Cloud and selecting available text-to-video, image-to-video, or storyboard workflows.
Aug 05, 2026 418 words in the original blog post.
MiniMax H3, a cutting-edge open-weights video model, is now supported in ComfyUI and offers a range of innovative features for video generation, including native stereo sound and 2K resolution outputs. This third-generation model from MiniMax allows for text-to-video, image-to-video, and reference-to-video capabilities, seamlessly integrating audio and visual elements to produce up to 15-second clips. Notably, it supports multimodal context understanding, allowing users to input text, images, audio, and video, which the model interprets to create cohesive outputs. The model's architectural advancements include significant memory optimization, enabling it to run efficiently on consumer-level hardware such as an RTX 3060 GPU. These enhancements are achieved by pruning modulation weights and utilizing int8 convrot quantization, reducing the memory footprint by 66%. Users can access MiniMax H3 through ComfyUI’s latest version, where they can download various workflows and templates to begin creating content.
Aug 03, 2026 1,589 words in the original blog post.