Home / Companies / Together AI / Blog / Post Details
Content Deep Dive

Improved Batch Inference API: Enhanced UI, Expanded Model Support, and 3000Ã Rate Limit Increase

Blog post from Together AI

Post Details
Company
Date Published
Author
Rajas Bansal, Mitali Meratwal, Nikitha Suryadevara, Will Van Eaton, Rishabh Bhargava
Word Count
374
Company Posts That Month
4
Language
English
Hacker News Points
-
Post removed?
No
Summary

The improved Batch Inference API offers significant enhancements, including a streamlined user interface, expanded support for all serverless models and private deployments, and a substantial increase in rate limits from 10 million to 30 billion enqueued tokens per model per user, representing a 3000× increase. This makes it simpler, faster, and more economical, operating at 50% of the cost of the real-time API for processing large-scale datasets. It is particularly advantageous for high-throughput tasks like large-scale text analysis, fraud detection, synthetic data generation, and content moderation, enabling teams like Inception Labs to conduct massive experiments efficiently. These updates aim to make large-scale inference more accessible and cost-effective, positioning the Batch Inference API as an ideal solution for handling extensive workloads without real-time constraints.

Trends Found in this Post
Trend Post Mentions Total Month Mentions Posts Companies MoM
Real-time 3 4,881 1,155 268 -10%
Serverless 2 961 189 88 +24%
AI Guardrails 1 428 112 48 +7%
Vector Search 1 1,772 362 150 +1%
Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.