May 2024 Summaries
8 posts from Vantage
Filter
Month:
Year:
Post Summaries
Back to Blog
Vantage has introduced a new feature that allows users to collect and report on Kubernetes GPU idle costs, enhancing their ability to identify underutilized resources within AI-intensive workloads. This feature is available to users with Vantage Kubernetes agent version 1.0.26 or later, and it requires the installation of the NVIDIA operator on their clusters. The new capability incorporates GPU memory usage into Kubernetes efficiency reports, which previously only calculated idle costs using CPU and RAM, thus providing a more comprehensive view of resource utilization. The data is gathered using the NVIDIA DCGM Exporter, and reports are updated within 48 hours as the costs from the infrastructure provider are ingested. This feature does not incur additional costs and supports whole GPU requests, with current compatibility limited to NVIDIA GPUs on AWS infrastructure.
May 30, 2024
926 words in the original blog post.
Vantage has introduced a new feature that allows users to upload labeled business metrics to calculate unit costs, enhancing the granularity of reporting for specific teams, categories, or business units. This update enables users to create a single business metric with a label parameter to segment data, avoiding the need for separate metrics for each application or category, thus simplifying dynamic cost allocation based on usage metrics. The labels, which can identify sources such as applications or cost centers, are available alongside existing date and amount parameters and can be used in Cost Reports to visualize unit costs across different categories. The feature supports all existing metric ingestion methods, including API, CSV files, Amazon CloudWatch, or Datadog, and is free for all users. The newly added label parameter is optional, allowing users to continue importing metrics without it, while still benefiting from improved decision-making capabilities at the application level. Users can manage and view these metrics and labels through the Business Metrics screen in the console, facilitating the assignment of specific labels to Cost Reports for detailed cost analysis.
May 29, 2024
1,094 words in the original blog post.
Vantage has launched Vantage University, a collection of training videos and guides designed to help users get started with its platform. Accessible through the Vantage product documentation site, this new educational resource offers video demonstrations, role-specific use cases, and an in-depth exploration of Vantage features, including cost reporting, budgeting, and cost allocation. Vantage University supplements the existing written materials, such as comprehensive product documentation, API guides, and the Cloud Cost Handbook, by providing visual content that caters to users who prefer learning through videos. The service is free of charge and aims to enhance the user experience by making it easier for both new and existing users to understand and utilize Vantage’s capabilities effectively. Feedback can be provided via the Slack Community or the product documentation GitHub repository, and additional content will be added as new features are developed.
May 28, 2024
712 words in the original blog post.
In the competitive landscape of small language models, Llama 3 8B and Mistral 7B offer distinct advantages and trade-offs for users seeking cost-effective AI solutions. Llama 3 8B, introduced by Meta in April 2024, features 8 billion parameters and excels in tasks like text summarization, sentiment analysis, and language translation, outperforming Mistral 7B on popular leaderboards. Despite its superior performance and broader language support, Llama 3 8B is more expensive, with its pricing through Amazon Bedrock being significantly higher than Mistral 7B, which is 62.5% less expensive for input tokens and 66.7% less for output tokens. Mistral 7B, a dense transformer model released in September 2023, remains popular for its balance between performance and affordability, making it a cost-efficient choice for tasks such as text summarization and code completion. While Llama 3 8B offers a faster inference speed and was trained on a more extensive dataset, Mistral 7B is more regionally available, being performative in English, and presents a compelling option for budget-conscious companies processing large volumes of data.
May 21, 2024
927 words in the original blog post.
A blog post by Emily Dunenfeld discusses the unexpected costs associated with Amazon S3, citing a case where a standard document indexing system incurred over $1,300 in charges due to unauthorized requests. This incident highlights the potential for S3 costs to escalate due to misconfigurations, security vulnerabilities, or third-party tool integrations, even when buckets are private or empty. AWS has made commitments to address some of these issues, but preventive measures like proactive monitoring and alerting are essential to avoid cost overruns. Tools such as Vantage offer comprehensive visibility into cloud spending, anomaly detection, and alerting features, enabling users to monitor costs by bucket, region, and API request type, among others. This approach helps organizations understand spending patterns, identify anomalies, and take corrective actions before costs spiral out of control.
May 14, 2024
1,187 words in the original blog post.
Vantage has introduced a new API endpoint for retrieving financial, optimization, and waste reduction recommendations, allowing users to programmatically access and filter recommendations by provider, account, and type. This enhancement builds upon the previous capability that only displayed recommendations in the user interface, offering more flexibility by enabling users to import recommendations into other systems. Additionally, the API allows users to access specific resources associated with recommendations, such as Kubernetes rightsizing. This feature is available for free to all users, including those on the free tier, and supports providers like AWS, Azure, Datadog, and Kubernetes, offering tailored suggestions such as EC2 rightsizing or Azure reserved instances. Authentication is managed via a Vantage API token, and users can refer to comprehensive API documentation and resources for further guidance.
May 13, 2024
804 words in the original blog post.
Vantage has enhanced its budgeting capabilities, allowing users to track budgets against forecasted costs over extended multi-month periods, a significant improvement from the previous system where budgets were only visible up to the date of accrued costs. This new feature enables users to visualize future budget periods alongside forecasted costs directly within a Cost Report, offering a more comprehensive view of both current and future fiscal performance. Users can select a forecast period through a date picker, allowing them to compare forecasted data with budgeted amounts for the same period. The feature is available to all Vantage users at no additional cost and supports various chart types and date binning options, although at present it does not support budget performance export. To further explore the budgeting functionalities, users are encouraged to consult the Budgets documentation available on the platform.
May 06, 2024
678 words in the original blog post.
The text discusses the growing popularity and demand for GPU instances in Amazon EC2, driven by advancements in AI, machine learning, gaming, and other compute-intensive applications. With the rise in these applications, companies are increasingly renting GPU compute power from cloud providers like Amazon to avoid the high upfront costs of hardware. Amazon offers two main categories of GPU instances: the P family, which is optimized for general-purpose GPU compute tasks and large-scale model training, and the G family, which is more cost-effective and suited for graphics-intensive applications and smaller ML workloads. The text compares these instances based on use cases, performance, instance size, and pricing, highlighting that while the P family is powerful and suited for demanding tasks, the G family strikes a balance between cost and performance, making it suitable for a wide range of workloads. It emphasizes the importance of considering factors like availability, hardware compatibility, and cost-effectiveness when choosing an appropriate instance for specific machine learning workloads.
May 02, 2024
1,518 words in the original blog post.