Home / Companies / Hugging Face / Blog / Post Details
Content Deep Dive

Async GRPO with LoRA across HF Jobs: a bucket, a proxy, and no NCCL

Blog post from Hugging Face

Post Details
Company
Date Published
Author
Amine Dirhoussi, Quentin Gallouédec, Kashif Rasul, and Sergio Paniego
Word Count
6,317
Company Posts That Month
82
Language
-
Hacker News Points
-
Post removed?
No
Summary

TRL v1.14 adds LoRA support to AsyncGRPOTrainer, enabling reinforcement learning training systems to synchronize only small LoRA adapters rather than full model weights with vLLM inference servers. The described implementation runs a trainer and multiple vLLM replicas as separate Hugging Face Jobs, using a shared Storage Bucket mounted as a filesystem to distribute versioned adapters without NCCL or direct inter-machine networking. A local proxy handles Hugging Face authentication, broadcasts adapter updates consistently across replicas, and routes rollout requests according to adapter-aware KV-cache prefix affinity while preventing load imbalance. Tests using Qwen2.5-Math-1.5B, rank-1 LoRA, and the Sanity dataset showed reliable policy consistency, with rollout ratios remaining near 1.0 throughout adapter updates. Performance investigations identified alternating training and generation bottlenecks, and improvements including token-budget batching, disabling unnecessary gradient checkpointing, reducing adapter-load retry intervals, and increasing request concurrency reduced a 500-step run from 3 hours 27 minutes to 53 minutes while maintaining similar reward progression.

Trends Found in this Post
Trend Post Mentions Total Month Mentions Posts Companies MoM
AI Model Fine-tuning 38 139 28 14 -75%
LLM 2 747 162 79 -85%
Real-time 1 649 155 80 -85%
Secrets Management 1 451 99 43 -80%
Serverless 1 156 54 28 -80%
Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.