August 2026 Summaries
1 posts from Fal
Filter
Month:
Year:
Post Summaries
Back to Blog
fal has introduced H3 Max, a post-trained and inference-optimized version of the open-weights MiniMax H3 video model designed to combine high visual quality, prompt adherence, and low latency. The company reports that H3 Max generates a five-second video in under three seconds, delivering roughly 35 times the throughput of the official MiniMax H3 endpoint while outperforming 12 competing video models in internal human-preference evaluations for overall quality, prompt understanding, and aesthetics. Developed through coordinated model post-training and systems optimization on NVIDIA GB200 NVL72 hardware, H3 Max retained only speed improvements that preserved its quality ranking, avoiding common performance tradeoffs such as reduced precision or fewer sampling steps when they harmed outputs. fal says independent benchmarks from Artificial Analysis and Design Arena also rank the model highly, and it positions H3 Max as suitable for interactive and high-volume production uses where quality, latency, and cost all matter. The model is available through fal’s Playground, Agent, and API, with an introductory discount and training infrastructure offered through fal Serverless.
Aug 27, 2026
980 words in the original blog post.