Home / Companies / Baseten / Blog / October 2024

October 2024 Summaries

3 posts from Baseten

Filter
Month: Year:
Post Summaries Back to Blog
With Baseten Hybrid, sensitive workloads can be run in a customer's own VPC with flex compute available on-demand on Baseten Cloud. Early access is now available for companies looking to scale their own VPC, offering special pricing until the end of October. This solution enables complete control over policies and workloads while gaining cloud elasticity, allowing seamless spillover from self-hosted workloads to Baseten Cloud with no engineering effort required. It combines the benefits of Baseten Self-hosted and Baseten Cloud, providing multi-cloud flexibility and true cloud agnosticism, with full support and management from a team of expert engineers. This solution is ideal for meeting specific security needs, achieving multi-cloud elasticity, spend down cloud commits, blazing-fast inference, and compound AI systems, while future-proofing infrastructure against traffic bursts without compromising on performance or security.
Oct 28, 2024 691 words in the original blog post.
The NVIDIA H200 Tensor Core GPU is designed for AI workloads and offers more GPU memory and memory bandwidth compared to its sibling, the popular H100 GPU. While it's anticipated for training, fine-tuning, and other long-running AI tasks, testing shows that H200 GPUs are a good choice for large models, large batch sizes, and long input sequences in terms of inference tasks. However, outside these situations, they offer minimal performance improvements over H100 GPUs, making them less cost-efficient for many inference tasks. The GH200 GPU may offer stronger inference performance in more circumstances.
Oct 23, 2024 1,294 words in the original blog post.
Baseten has introduced an export metrics integration that allows users to export model inference metrics like response time, replica count, and hardware utilization to observability platforms such as Grafana, New Relic, Datadog, and Prometheus. This feature enhances production model management workflows by providing a single source of truth, fine-grained metrics and control, and custom alerts. The supported metrics include inference request count, end-to-end response time, replica count, and hardware usage for CPU and GPU resources. Metrics are exported using the vendor-neutral OpenTelemetry standard and can be scraped by observability tools at a set interval. Baseten's integration supports over 800 tools in the OpenTelemetry registry, with specific documentation available for popular platforms like Grafana, New Relic, Datadog, and Prometheus.
Oct 05, 2024 493 words in the original blog post.