ERNIE-Image Day-0 Support in ComfyUI: Precise Text Rendering and Structured Image Generation
Blog post from Comfy
ERNIE-Image, Baidu's open-source text-to-image model licensed under Apache-2.0, is now available on ComfyUI, utilizing an 8 billion parameter Diffusion Transformer to deliver precise text rendering, strong instruction following, and structured visual generation across a wide stylistic range. It features a built-in Prompt Enhancer, which uses a 3 billion parameter model to expand short inputs into richer prompts, enabling the creation of complex and detailed artworks, from educational infographics to cinematic posters and editorial fashion photography. The model supports dense and layout-sensitive text in multiple languages and is compact enough to run on 24 GB VRAM, allowing for both realistic and stylized image generation. Users can download and explore different versions such as the Main SFT model for high-quality output and the ERNIE-Image-Turbo for faster generation, both easily accessible through Comfy Cloud, enhancing creative workflows for artists and developers alike.
No tracked trend matches for this post yet.
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.