Home / Companies / Anyscale / Blog / Post Details
Content Deep Dive

Streaming distributed execution across CPUs and GPUs

Blog post from Anyscale

Post Details
Company
Date Published
Author
Eric Liang, Stephanie Wang, Cheng Su
Word Count
2,067
Company Posts That Month
9
Language
English
Hacker News Points
-
Post removed?
No
Summary

Ray Data provides streaming execution for large-scale batch inference workloads, offering improved performance on heterogeneous clusters with both CPU and GPU devices. This allows for pipelined execution across an entire cluster, avoiding unnecessary overheads associated with bulk synchronous parallel frameworks. By leveraging end-to-end pipelining, Ray Data can handle demanding use cases such as video decoding, annotation, and classification, while also providing optimizations like memory stability, data locality, and fault tolerance to ensure seamless execution. The streaming backend is fully backwards compatible with the existing API, allowing users to transform datasets lazily with map operations and support shuffle operations, caching / materialization in memory, and more. Early users are taking advantage of Ray Data streaming to create efficient large-scale inference pipelines over unstructured data, including video and audio data.

Trends Found in this Post
Trend Post Mentions Total Month Mentions Posts Companies MoM
Real-time 24 1,875 540 158 +10%
Observability 1 1,402 256 72 +41%
Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.