Home / Companies / Anyscale / Blog / Post Details
Content Deep Dive

Ray Serve LLM on Anyscale: APIs for Wide-EP and Disaggregated Serving with vLLM

Blog post from Anyscale

Post Details
Company
Date Published
Author
Seiji Eicher
Word Count
1,281
Company Posts That Month
12
Language
English
Hacker News Points
-
Post removed?
No
Summary

Ray Serve LLM introduces new APIs that facilitate the deployment of advanced serving patterns for sparse mixture-of-experts models, like DeepSeek and Qwen3, using vLLM on the Anyscale platform. The APIs support wide expert parallelism and disaggregated prefill/decode serving, enabling models to achieve high throughput and optimized latency by balancing expert loads and separating processes that handle input prompts from those generating output tokens. By leveraging Ray Serve, developers can build and orchestrate complex model deployments using Pythonic builder patterns, which allow for dynamic scaling, stateful routing, and fault-tolerant orchestration, while maintaining compatibility with Kubernetes environments. This approach reduces the operational burden of coordinating multi-node setups and enhances performance through programmable orchestration, enabling efficient use of resources and maintaining high service level agreements.

Trends Found in this Post
Trend Post Mentions Total Month Mentions Posts Companies MoM
LLM 22 5,556 752 184 +14%
Kubernetes 3 1,297 225 80 -9%
Observability 1 2,534 521 146 +9%
Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.