How Notion cuts embedding costs by 80% and other stories on scaling AI with Ray from Salesforce, Uber, and more…
Blog post from Anyscale
Ray Day Seattle, part of Anyscale's 2026 Ray on the Road series, showcased how companies like Notion, Salesforce, Uber, and Apple are leveraging the Ray framework to scale AI efficiently. At the event, Notion explained its migration from a Spark-based embedding pipeline to a streamlined Ray-powered job, achieving an 80% cost reduction and significant improvements in query latency. Salesforce discussed its document summarization pipeline using Ray, which processes up to 200K tokens with a P95 latency under 15 seconds, highlighting Ray's role in parallelizing tasks. Uber's presentation focused on enhancing GPU utilization and reducing training time by adopting Ray, while Apple demonstrated Ray's capability in handling massive foundation model training by unifying data processing and model training. The day also featured workshops for developers to deepen their understanding of Ray's core functionalities and its application in scalable data pipelines, distributed training, and production model serving.
| Trend | Post Mentions | Total Month Mentions | Posts | Companies | MoM |
|---|---|---|---|---|---|
| Vector Search | 8 | 1,739 | 413 | 146 | -27% |
| LLM | 4 | 5,932 | 1,046 | 223 | -2% |
| Observability | 1 | 4,496 | 812 | 176 | +40% |
| RAG | 1 | 941 | 216 | 85 | -48% |
| Reinforcement learning | 1 | 104 | 49 | 23 | -14% |
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.