New in February 2024
Blog post from Baseten
NVIDIA has released new improvements in February 2024, focusing on model performance across four key factors: latency, throughput, quality, and cost. The company now offers model inference on H100 GPUs, which feature exceptional performance for running ML models due to their high tensor compute, memory bandwidth, and VRAM. This results in a significant reduction in cost for running high-traffic workloads. Additionally, NVIDIA has optimized Stable Diffusion XL with TensorRT, achieving 40% lower latency and 70% higher throughput on H100 GPUs compared to a baseline implementation. The company has also introduced SDXL Lightning, which generates images in under one second per image, while QwenVL is an open-source visual language model that combines vision and language capabilities. Furthermore, NVIDIA's refreshed billing dashboard provides daily insights into usage and spend, offering improved visibility for users.
| Trend | Post Mentions | Total Month Mentions | Posts | Companies | MoM |
|---|---|---|---|---|---|
| Vector Search | 5 | 2,087 | 216 | 81 | +23% |
| LLM | 4 | 2,401 | 292 | 122 | -7% |
| Real-time | 1 | 2,379 | 618 | 172 | -8% |
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.