Home / Companies / Together AI / Blog / Post Details
Content Deep Dive

Introducing Together AI Chief Scientist Tri Dao, as he releases FlashAttention-2 to speed up model training and inference

Blog post from Together AI

Post Details
Company
Date Published
Author
Together
Word Count
2,001
Company Posts That Month
5
Language
English
Hacker News Points
-
Post removed?
No
Summary

Tri Dao, a Chief Scientist at Together AI, has released FlashAttention-2, an algorithm designed to speed up training and inference of large language models by up to 4x and achieving 72% model FLOPs utilization on NVIDIA A100 GPUs. The new version is built from scratch using primitives from NVIDIA's CUTLASS 3.x and its core library CuTe, providing clean abstractions and powerful building blocks for maximum speed. FlashAttention-2 achieves a 2x speedup over the previous implementation, reaching up to 230 TFLOPs/s on A100 GPUs, and is available in open source on Github. The algorithm is designed to work with existing models and can be used for training, fine-tuning, and inference of large language models. With its improvements, FlashAttention-2 enables models with twice as long context lengths while maintaining an interactive experience, making it a significant breakthrough in the field of natural language processing.

Trends Found in this Post
Trend Post Mentions Total Month Mentions Posts Companies MoM
AI Model Fine-tuning 5 669 87 53 +50%
LLM 4 1,935 244 98 -1%
Real-time 1 2,035 534 182 -15%
Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.