Supercharging NVIDIA H200 and H100 GPU Cluster Performance With Together Kernel Collection
Blog post from Together AI
The NVIDIA H200 Tensor Core GPU is a high-performance computing (HPC) and artificial intelligence (AI) workhorse, designed to excel in both AI and HPC workloads. With its advanced Hopper architecture, the H200 provides 40% faster inference performance on Llama 2 13B and 90% faster performance on Llama 2 70B, demonstrating significant improvement in handling large-scale language models. The GPU's substantial memory and bandwidth allow it to handle even the most data-intensive applications with ease, minimizing bottlenecks and enabling real-time processing of vast datasets. Together AI's custom-built Together Kernel Collection (TKC) offers up to 24% speedup for operators used frequently in training and up to 75% speedup for fundamental operations used in FP8 inference, significantly accelerating common AI operations.
| Trend | Post Mentions | Total Month Mentions | Posts | Companies | MoM |
|---|---|---|---|---|---|
| Kubernetes | 4 | 1,327 | 196 | 88 | +0% |
| LLM | 4 | 4,030 | 486 | 147 | +1% |
| AI Model Fine-tuning | 3 | 685 | 161 | 75 | -31% |
| Observability | 1 | 1,798 | 331 | 106 | +34% |
| Real-time | 1 | 4,377 | 976 | 225 | +49% |
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.