Home / Companies / Anyscale / Blog / Post Details
Content Deep Dive

Announcing Native LLM APIs in Ray Data and Ray Serve

Blog post from Anyscale

Post Details
Company
Date Published
Author
The Anyscale Team
Word Count
1,038
Company Posts That Month
2
Language
English
Hacker News Points
-
Post removed?
No
Summary

Ray Data LLM provides APIs for offline batch inference with LLMs within existing Ray Data pipelines, while Ray Serve LLM offers APIs for deploying LLMs for online inference in Ray Serve applications. Both modules offer first-class integration for vLLM and OpenAI compatible endpoints, addressing common developer pains around batch inference, such as launching multiple online inference servers and proxying/load balancing utilities to maximize throughput. Ray Data LLM simplifies the usage of LLMs within existing data pipelines by providing a Processor object that can be called on a Ray Data Dataset, while Ray Serve LLM allows users to deploy multiple LLM models together with a familiar Ray Serve API, offering features like automatic scaling and load balancing, unified multi-node multi-model deployment, OpenAI compatibility, and composable multi-model LLM pipelines.

Trends Found in this Post
Trend Post Mentions Total Month Mentions Posts Companies MoM
LLM 31 4,226 639 179 -13%
Serverless 2 1,599 300 96 +114%
AI Model Fine-tuning 1 697 168 71 +1%
Kubernetes 1 2,271 264 89 +53%
Real-time 1 6,887 1,132 212 +49%
Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.