Batch API: half-price inference by bundling requests
Blog post from OpenRouter
OpenRouter has introduced an asynchronous Batch API for non-urgent, high-volume workloads such as corpus labeling, embedding backfills, evaluation scoring, ticket summarization, and large-scale prompt execution, generally offering discounts of 50% or more from normal per-token pricing across more than 70 models. Requests can use chat completions, responses, messages, or embeddings formats and are submitted to a batch endpoint, then polled until they are completed, failed, expired, or cancelled, with completed results returned inline. Although providers have up to 24 hours to process work, beta data from more than 230,000 batches showed a median completion time of seven minutes, with 90% completed within an hour and 99% within 10.3 hours; submission time affected speed more than request volume, with batches submitted between 5 a.m. and noon Pacific taking longer. Each batch runs on one provider selected according to pricing and user routing settings, supports eligible bring-your-own-key configurations, returns results independently to prevent isolated failures from stopping an entire job, and retains inputs and outputs for 30 days unless deleted. Batch pricing varies by model, while web search remains normally priced, and limitations include requiring public URLs for image and file inputs and excluding audio, video, and OpenRouter’s web search plugin.
| Trend | Post Mentions | Total Month Mentions | Posts | Companies | MoM |
|---|---|---|---|---|---|
| Vector Search | 3 | 265 | 57 | 33 | -89% |
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.