August 2026 Summaries
15 posts from Anyscale
Filter
Month:
Year:
Post Summaries
Back to Blog
Learning loops are presented as a way for companies to build durable AI advantages by continuously connecting proprietary data collection and curation, custom model training, evaluation, and inference so that each cycle improves performance, efficiency, or cost. The proposed maturity path begins with using rented models while developing proprietary prompts and evaluations, progresses to owning model weights and runtime infrastructure when cost, reliability, security, and data sovereignty become important, and culminates in optimizing the entire loop around measures such as cost per unit of work rather than cost per token. Building these systems requires managing large multimodal datasets, long-running distributed GPU training, and elastic low-latency inference across varied hardware and cloud environments. The described architecture combines Ray for distributed, fault-tolerant workload scheduling, Kubernetes-oriented tools such as Kueue, Kai, Volcano, and Yunikorn for job prioritization, and a global control plane for visibility, access, placement, and resource coordination, with Anyscale and Ray positioned as tools intended to support this production AI infrastructure.
Aug 25, 2026
1,158 words in the original blog post.
Ray Core improvements target scaling bottlenecks in data pipelines and large AI training workloads, where millions of objects, tens of thousands of actors, and clusters up to 10,000 nodes exposed lock contention, overloaded control threads, and scheduling based on stale resource information. For Ray Data batch inference and shuffle, changes such as moving publishing work to a background thread, reducing unnecessary callbacks and object-location subscriptions, and optimizing local-object handling increased sustained task throughput by up to 1.6×, producing 23% faster batch inference and 24% faster shuffle performance on 500 nodes. For large reinforcement-learning and training clusters, Ray made its Global Control Service multithreaded, reduced control-plane work, removed an unnecessary scheduling reset for placement groups, and replaced dummy readiness tasks with direct asynchronous queries, enabling placement groups to become ready 62× faster at 2,000 nodes and 303× faster at 10,000 nodes. The updated system can support 40,000 actors on a 10,000-node cluster and launches actors 6.5× faster at 2,000 nodes, while ongoing work seeks to centralize more scheduling decisions and move non-detached actor lifecycle management closer to their owners to further reduce startup and recovery delays.
Aug 25, 2026
3,830 words in the original blog post.
Ray Data’s Shuffle V2 redesigns how key-based operations such as joins, groupbys, repartitions, and deduplication handle intermediate data, replacing V1’s long-lived aggregator actors with separate map and reduce stages connected through materialized shards in the Ray object store. This approach allows intermediates to spill to disk, use Ray’s lineage-based reconstruction after failures, measure actual partition sizes to reduce out-of-memory risks, and release compute resources between stages rather than reserving them for an entire shuffle. V2 also introduces input coalescing to reduce object counts, configurable Arrow IPC compression, downstream map fusion, and Arrow-native vectorized aggregations for several common functions. Benchmarks on TPC-H SF1000 workloads report major gains for aggregation-heavy queries and successful completion of joins that V1 could not finish, although V2’s architecture alone can be slower for small shuffles due to encoding and object-store overhead. Available experimentally in Ray 2.58 as `HASH_SHUFFLE_V2` and renamed `SHUFFLE_V2` in Ray 2.59, the engine is intended to eventually support sort and random shuffle workloads, while planned disk shuffle and incremental join features aim to further reduce memory and object-management limits.
Aug 25, 2026
2,593 words in the original blog post.
Anyscale has announced a private preview of GPU Health Observability, a tool designed to connect GPU hardware telemetry directly to the Ray jobs and workspaces affected by it, helping teams distinguish hardware faults from application or training-code failures. Compatible with KubeRay and virtual machines, with planned support for Kubernetes environments managed by the Anyscale Operator, the feature uses DCGM metrics such as XID errors, ECC memory errors, clock speeds, memory use, temperature, power, and NVLink indicators, enriched with workload, node, and GPU context. Platform operators can use a fleet-level console view to identify unhealthy nodes and inspect individual GPUs, while ML engineers can see hardware-error alerts within job views that identify the affected node, GPU, and workers. By consolidating infrastructure and runtime information into a correlated interface, Anyscale aims to reduce manual investigation and mean time to resolution for training and inference issues. The company plans to expand the observability layer with more granular health analysis, fault tolerance capabilities, and broader Kubernetes support.
Aug 25, 2026
1,804 words in the original blog post.
Anyscale KubeRay Connect, introduced in private preview, is designed to let organizations using Ray on Kubernetes adopt Anyscale observability and orchestration without replacing their existing KubeRay-based workflows, Custom Resources, or Kubernetes tools such as Kueue, Argo CD, and Kyverno. Rather than directly managing Pods, it integrates with KubeRay CRDs through a Helm-installed Connector that mutates workload specifications to add telemetry components while leaving reconciliation to the KubeRay Operator. Its initial Observe tier provides dashboards, logs, events, metrics, and Telescope full-stack visibility across hardware, Kubernetes infrastructure, and Ray applications, while the private-preview Orchestrate tier supports workload submission, status tracking, and scheduling capabilities; an Optimize tier for price-performance tuning is planned. The architecture separates control and telemetry functions to support resilience, independent scaling, egress-only networking, and Kubernetes-native authentication, with workloads continuing to run during Anyscale control-plane outages. Future development includes multi-cluster orchestration, identity passthrough for access controls and quotas, and multi-region serving support.
Aug 25, 2026
1,360 words in the original blog post.
Ray Data 2.58 introduces GPU-native processing capabilities through cuDF support for GPU-accelerated DataFrame batches and an experimental RapidsMPF-based shuffle backend for GPU repartitioning and hash aggregations. Previously, Ray Data could schedule GPU-equipped user-defined functions but relied on users to manually invoke GPU libraries; the new features make GPU execution more integrated for AI data workloads such as multimodal processing and training-data preparation. Benchmarks of fuzzy document deduplication on the FineWeb 10BT dataset found that cuDF accelerated MinHash signature generation by roughly 15 times over comparable CPU clusters, while GPU shuffle improved grouping and connected-components stages by up to 3 times, although communication overhead reduced gains at larger GPU counts. Across the full pipeline, the GPU implementation was up to four times faster and achieved up to 3.1 times better total cost of ownership than CPU alternatives. Future work includes operator fusion to keep data in GPU memory between stages and GPU shuffle support for Ray Data preprocessors.
Aug 25, 2026
1,430 words in the original blog post.
KubeRay v1.6 and v1.7 introduce a broad set of improvements for operating Ray workloads on Kubernetes, incorporating more than 381 commits from roughly 75 contributors. Major additions include the beta KubeRay History Server, which preserves logs, events, Ray Data datasets, and Ray Serve application information after clusters are deleted, alongside automated collector sidecar injection and compatibility with token-authenticated clusters. Security enhancements add native NetworkPolicy configuration, cert-manager-based mTLS certificate automation, and Kubernetes RBAC-based token authentication, while RayService incremental upgrades gain beta status with rollback support. The releases also add cron-based RayJob scheduling, independently restartable job-submitter sidecars, improved job cleanup policies, a RayCluster recreate upgrade strategy, configurable Ingress settings, Autoscaler V2 worker-group priorities, and beta stable indexing for multi-host workers. For fault tolerance, KubeRay supports an embedded RocksDB backend for Ray GCS to reduce dependence on external Redis, improves Redis cleanup handling during deletion, and surfaces Kubernetes platform events in the Ray Dashboard to aid infrastructure troubleshooting.
Aug 25, 2026
3,190 words in the original blog post.
SkyRL has added FP8 acceleration for reinforcement-learning training and rollout, including FP8 linear-layer computation, model weights, KV caches, and trainer-to-vLLM weight transfers, while retaining BF16 or FP32 for precision-sensitive operations and optimizer state. The work addresses a key RL reliability issue in which independently quantizing weights in the trainer and rollout engines creates numerically different policies, leading to unstable training behavior; SkyRL’s on-policy weight synchronization instead transfers the trainer’s FP8 weight payloads and scale metadata directly so rollout weights are bitwise identical after synchronization. In 400-step DAPO experiments with Qwen3.5-9B on eight H100 GPUs and Qwen3.5-35B-A3B on eight B200 GPUs, FP8 with this synchronization method closely matched BF16 reward, pass@8, and response-length trends while reducing maximum end-to-end step time by about 19% and 23%, respectively, largely through faster memory-bound generation. FP8 parameter storage also reduced per-GPU trainer weight memory by 39–42% on Hopper systems, although training-phase gains were limited by scale-management and host-dispatch overhead, and support varies by hardware and quantization recipe.
Aug 25, 2026
3,097 words in the original blog post.
Ray 2.58 introduces experimental native sandboxing for agentic reinforcement learning and other workloads that execute model-generated code at scale, integrating gVisor-based isolated OCI container environments directly into Ray’s scheduling, resource management, autoscaling, and fault-tolerance systems. Users can employ a high-level API to create and manage sandbox actors across a cluster or use lower-level runtime primitives to build custom sandbox services, with controls for CPU, memory, networking, file transfers, environment configuration, privileges, and OCI specifications. Ray reports scaling to 100,000 gVisor sandboxes in 20 seconds on Google Kubernetes Engine, positioning sandbox placement as another distributed scheduling task rather than requiring a separate control plane. The feature supports practical patterns such as configurable network isolation, cross-node artifact transfer, MCP-based tools for safe code execution, and stricter security controls including capability removal and process limits. It also integrates with Harbor for coding-agent benchmark evaluations, although some task types, including Compose, GPU, and allowlist-based workloads, are not yet supported. Planned improvements include a REST service, GPU-backed sandboxes, Docker support within sandboxes, expanded network and filesystem functions, and security configuration guidance.
Aug 25, 2026
2,254 words in the original blog post.
Ray History Server, promoted to beta in KubeRay v1.7, provides post-mortem observability for ephemeral Ray clusters on Kubernetes by preserving and reconstructing Ray Dashboard data after clusters terminate. It offers a centralized interface for discovering, filtering, and opening both live and historical sessions, proxying live dashboards while rebuilding terminated ones from logs, structured events, and endpoint snapshots stored in object storage. A collector sidecar on Ray head and worker pods records telemetry using disk-first event collection, log uploads, metadata capture, and safeguards for disk pressure, while the stateless history server lazily retrieves and decompresses session data and uses an LRU cache to limit memory consumption. The storage-agnostic design supports Google Cloud Storage, Amazon S3 and MinIO, Alibaba Cloud OSS, and Azure Blob Storage, with deterministic paths that organize data by cluster and workload ownership. Benchmarks indicate that gzip compression and at least 1.5 CPU cores can reduce event storage by about 91%, speed cold session loads by roughly 2.4 times, and allow sessions of up to 50,000 tasks to open within a practical UI time budget. Planned improvements include automated collector injection through the KubeRay operator and replayable historical metrics such as CPU, GPU, memory, and actor-throughput data.
Aug 25, 2026
1,851 words in the original blog post.
Large language model serving requires routing requests across replicas while accounting for stateful KV caches, highly variable input and output lengths, and unpredictable generation costs, making conventional microservice load balancing insufficient. Ray Serve LLM offers session affinity, prefix affinity, and KV cache affinity approaches to reuse cached computation, but the post argues that maximizing cache reuse alone can create request herding and uneven workloads when long or high-token requests concentrate on the same replica. Its KVAwareRouter combines actual KV cache overlap, informed by vLLM cache events and NVIDIA Dynamo’s KV indexer, with token-load estimates that capture uncached prefill work and active decode demands. Tests using asynchronous reinforcement-learning rollouts and reconstructed Claude Code agent traces found that this approach may sacrifice some cache-hit rate but improves balance, time to first token, time per output token, throughput, and tail latency for heterogeneous workloads. Consistent-hash session affinity remains useful when cache locality is especially valuable and sessions have comparable workloads, while future work aims to address unknown output lengths and expand token-aware routing to disaggregated, data-parallel, and multimodal deployments.
Aug 25, 2026
2,655 words in the original blog post.
CVE-2025-62593 affects Ray releases before version 2.52.0 and allows attackers to exploit a bypassable User-Agent check in the dashboard and job submission API, potentially executing shell commands on a developer’s local machine through a malicious webpage combined with DNS rebinding. Ray fixed the issue in version 2.52.0, released November 26, 2025, by implementing proper browser-origin controls, and CISA added the vulnerability to its Known Exploited Vulnerabilities catalog on August 17, 2026, citing active exploitation of unpatched systems. Users running Ray 2.52.0 or later are not affected, while those on earlier versions should upgrade, enable the opt-in token authentication introduced in 2.52.0, keep dashboards within tightly controlled network boundaries, avoid unnecessary binding to all interfaces, and check pinned dependencies, images, CI environments, and long-lived development systems for outdated versions.
Aug 19, 2026
950 words in the original blog post.
Ray Serve’s asynchronous inference capability is presented through a video-indexing service that immediately accepts an S3 video URI, queues the job, and processes it in the background by downloading and chunking video frames with FFmpeg, embedding them with SigLIP on GPUs, and storing vectors in S3. The approach separates short client requests from long-running work through message queues, task retries, dead-letter handling, polling, and queue-depth-based autoscaling, helping services absorb bursts without maintaining long-lived HTTP connections. In a 20-minute test generating roughly 67,000 requests at sustained overload, the Ray Serve deployment scaled from one to four GPU replicas, later returned to one replica, and reported no failed or lost requests. A comparison using the same four NVIDIA T4 GPUs, fused single-worker architecture, workload, and cold start against Amazon SageMaker Async Inference found that both systems processed all requests, while Ray Serve reportedly reached full capacity faster, released idle capacity much sooner, and required a simpler autoscaling configuration; the post attributes much of SageMaker’s slower response to its CloudWatch metric evaluation interval. The article concludes that asynchronous inference is particularly suited to long-running, bursty machine-learning workloads and notes that Ray Serve can be deployed across cloud, on-premises, and local environments.
Aug 18, 2026
2,256 words in the original blog post.
Ray Direct Transport (RDT) is a Ray Core feature that simplifies RDMA-backed GPU-to-GPU tensor transfers between actors, targeting fast model-weight synchronization for reinforcement learning workflows involving large language models. RDMA bypasses operating-system networking layers to reduce latency and CPU overhead, but its benefits can be undermined by costly memory registration, metadata exchange, staging buffers, and inefficient transfers of many small tensors. Using a trainer and inference-generator example, the post explains how RDT integrates with NIXL to manage these operations while retaining Ray’s actor-based programming model. It recommends preregistering persistent tensor memory with `register_nixl_memory`, receiving data directly into model-weight buffers through `set_target_for_ref` to avoid duplicate memory use, and using preregistered memory pools to bucket small tensors into contiguous transfers. On two GB200 nodes connected through Multi-Node NVLink, these techniques achieved up to 859 GB/s for data transfer and improved end-to-end synchronization performance by as much as 7.5 times over a naive implementation. The authors also describe planned and ongoing integrations with RL frameworks including SkyRL and Miles, alongside further work to reduce metadata and Python-object transfer overheads.
Aug 18, 2026
3,444 words in the original blog post.
Ray has introduced NVLink Domain-Aware Placement Groups to improve scheduling on NVIDIA GB200 and GB300 NVL72 rack-scale systems, which connect 72 Blackwell GPUs and 36 Grace CPUs through a high-bandwidth NVLink fabric. The feature enables users to require related Ray actors and resource bundles to be colocated within one NVLink Domain, typically a rack, rather than being distributed across nodes or racks as with earlier node-focused placement policies. This topology awareness can improve collective communication performance by keeping GPU-intensive operations such as all-reduce on NVLink, while also preserving placement intent during node failures or maintenance by seeking replacement capacity within the same domain. In NVIDIA GEAR’s 512-GPU vision-language-action training evaluation, domain-aware placement organized 64-GPU groups within individual NVLink Domains and delivered 1.13 times faster iterations than a less localized deployment. The approach may also benefit disaggregated inference, where prefill and decode workers exchange cache data, and reinforcement learning pipelines with tightly coupled training components. Planned extensions include additional placement strategies such as spreading workloads across racks, support for hierarchical infrastructure topologies, and improved visibility into GPU locality.
Aug 13, 2026
1,603 words in the original blog post.