Building an RL theorem-proving workflow on Modal
Blog post from Modal
AE Studio describes using Modal to train language models for Lean theorem proving with reinforcement learning, comparing Evolution Strategies (ES), which evaluates randomly perturbed model variants and updates toward higher-scoring ones, with the more common Group Relative Policy Optimization (GRPO). The system separates GPU-based proof generation, CPU-based Lean verification, and orchestration, using Modal’s function-specific environments, parallel job mapping, isolated sandboxes, shared model volumes, and secrets management to coordinate thousands of proof attempts while containing verifier failures. ES checkpointing stores only deterministic perturbation seeds and rewards, allowing workers to reconstruct model updates without transferring large weight files. AE Studio reports that its implementation reduced infrastructure code and setup time, estimated lower costs through elastic GPU use, and produced early results in which ES sometimes matched or exceeded GRPO in verified proofs, particularly under limited training data, though performance varied and requires further study. The team plans experiments on hyperparameters, larger models, and ES scaling behavior, while presenting the architecture as applicable to other workflows that combine GPU generation with external verification or testing.
| Trend | Post Mentions | Total Month Mentions | Posts | Companies | MoM |
|---|---|---|---|---|---|
| LLM | 5 | 6,889 | 1,263 | 265 | -9% |
| Reinforcement learning | 4 | 109 | 54 | 27 | -40% |
| AI Model Fine-tuning | 2 | 472 | 158 | 73 | -60% |
| Secrets Management | 1 | 1,971 | 393 | 127 | +1% |
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.