Home / Trends / Reinforcement learning

Reinforcement learning

Reinforcement and preference learning for AI models

Explore historical trend data

Sign up for free to navigate to earlier time periods.

Tracked keywords: reinforcement learning, reward-based learning, rlhf, agent learning via feedback, preference learning, fine-tuning loops
Add filters to this data

Mentions Over Time

2,986
Total Mentions
1,225
Posts with Topic
19
Companies Mentioning
-10.2%
MoM Change

Historical Data

Month Mentions Posts Companies MoM Change
Aug 2026 88 39 19 -10.2%
Jul 2026 98 52 31 +24.1%
Jun 2026 79 44 27 -20.2%
May 2026 99 49 28 -2.0%
Apr 2026 101 54 27 -42.0%
Mar 2026 174 74 42 +31.8%
Feb 2026 132 62 39 -6.4%
Jan 2026 141 55 30 +34.3%
Dec 2025 105 56 32 -6.3%
Nov 2025 112 58 32 +5.7%
Oct 2025 106 41 22 +9.3%
Sep 2025 97 33 24 -6.7%
Aug 2025 104 48 32 -37.0%
Jul 2025 165 64 36 +68.4%
Jun 2025 98 48 32 -13.3%
May 2025 113 93 31 -46.9%
Apr 2025 213 96 26 +8.1%
Mar 2025 197 61 40 +9.4%
Feb 2025 180 54 33 +5.3%
Jan 2025 171 40 24 +288.6%
Dec 2024 44 29 17 +29.4%
Nov 2024 34 20 16 -48.5%
Oct 2024 66 20 18 -75.5%
Sep 2024 269 35 15 +389.1%

Recent Blog Posts

50 posts
Date Company Title Mentions
2026-09-03 Hugging Face Training a coding model to paint watercolours with TRL and OpenEnv 1
2026-09-02 Baseten Best open-source models for post-training 5
2026-08-31 Fireworks AI Train past the frontier: Training API now generally available 1
2026-08-28 Prime Intellect GLM-5.2 RL weight transfer in 4 seconds using NIXL and ModelExpress 1
2026-08-27 AssemblyAI How to learn machine learning in 2026: an updated roadmap 1
2026-08-27 Hugging Face TAVR: Generate Your Talking Avatar from Video Reference 1
2026-08-27 JetBrains Differential Privacy for Hugging Face Trainers – Without Rewriting Your Trainin… 2
2026-08-26 Fireworks AI Post-training Kimi K3 with Harvey for long-horizon legal work 2
2026-08-25 JetBrains Ideas Worth a Longer Conversation: The JetBrains Research Podcast - The JetBrai… 1
2026-08-25 LaunchDarkly ML Experiment Tracking: What to Track Across Models, Data, and Production 1
2026-08-25 LaunchDarkly Best Practices for Experiment Tracking in MLOps 2
2026-08-25 Anyscale FP8 Reinforcement Learning in SkyRL: Preserving Policy Consistency Across Train… 5
2026-08-25 Hugging Face Granite 4.2 LLMs: How They're Built 12
2026-08-25 Hugging Face Quantization-Aware Healing: a compressed, 4-bit model that outperforms its full… 1
2026-08-25 Lambda AgentFlow: when the agent's workflow learns 2
2026-08-24 Voxel51 AV playbook for robotics: what transfers and what doesn't 1
2026-08-24 Encord Teleoperation vs. Simulation: Where Should Your Robot Training Data Actually Co… 1
2026-08-23 LaunchDarkly ML Experiment Tracking: What to Track Across Models, Data, and Production 2
2026-08-19 Anyscale CVE-2025-62593 and the CISA KEV listing: what Ray users need to know 1
2026-08-19 Encord The Complete Physical AI Data Pipeline: From Data Collection to Deployment 1
2026-08-18 Anyscale Async inference in practice: a video-indexing service on Ray Serve 1
2026-08-18 Anyscale Using Ray Direct Transport for Fast and Easy Weight Syncing in Reinforcement Le… 4
2026-08-18 Redis ReAct agents explained: concepts & practical uses 1
2026-08-17 Crowdstrike Teaching AI to Reason Through Detection Triage 2
2026-08-17 Hugging Face Same Cluster, 33 Points More Utilization: What Changed Was the Order 2
2026-08-17 Lambda A world model for market microstructure 1
2026-08-14 Eden AI GLM-5.3 Benchmark vs GPT-5.6 Sol, Claude Fable 5 & Gemini 3.1 Pro 1
2026-08-14 RunPod The six AI model families and what they're good for 2
2026-08-13 Eden AI Google DeepMind Reshuffle: Hassabis Moves to Chair as Jeff Dean Departs 2
2026-08-13 Hugging Face Sleeper Agents and How to Tame Them 2
2026-08-13 Anyscale Maximizing the Power of NVIDIA GB300 NVL72: NVLink Domain-Aware Placement Group… 1
2026-08-13 Arcee AI Introducing NAC, an Open-Source Harness for Long-Running Agent Work 1
2026-08-12 Patronus AI Getting GLM-5.2 NVFP4 Post-Training off the ground 1
2026-08-12 Hugging Face LFM2.5-VL-3B for Better and Faster Vision Capabilities for the Edge 1
2026-08-11 Baseten Introducing NVIDIA Nemotron 3.5 Lightning 1
2026-08-11 CodeRabbit Teaching NVIDIA Nemotron 3.5 Lightning to route code reviews 1
2026-08-11 Hugging Face Luth-2: Pushing the French Capabilities of SLMs with MOPD 4
2026-08-07 Tavily The Rise of Enterprise Learning Sovereignty 8
2026-08-07 TestMu AI What is AI Model Testing: Methods & Best Practices 1
2026-08-06 Hugging Face FP8 KV-Cache on Intel® Arc™ Pro B70: 2× Capacity with strong Long Context Throu… 1
2026-08-05 AssemblyAI How to build real-time agent assist on streaming speech-to-text 2
2026-08-04 Hugging Face Deploy local agents everywhere with LFM2.5-2.6B 2
2026-08-03 Hugging Face 超越表层对齐:信念是通往深层对齐的新入口 10
2026-07-31 Eden AI Moonshot Distilled Anthropic's Fable: What Model IP Theft and Treasury Sanction… 4
2026-07-31 Eden AI Tiny LLMs That Beat Giant Models: How Efficient 3B-Parameter Models Compete wit… 2
2026-07-31 n8n How LLM Guardrails Keep Production AI Safe 2
2026-07-30 Anyscale Anyscale signs definitive agreement to join Nscale 1
2026-07-30 Arcee AI Teaching an Open Model to Do Science 1
2026-07-29 Braintrust Behavior specs, an open standard for supervising long-horizon agents 1
2026-07-29 Hugging Face VisionPsy-Nano: State-of-the-Art On-Device Vision-Language Models 2