Home / Trends / Reinforcement learning

Reinforcement learning

Reinforcement and preference learning for AI models

Explore historical trend data

Sign up for free to navigate to earlier time periods.

Tracked keywords: reinforcement learning, reward-based learning, rlhf, agent learning via feedback, preference learning, fine-tuning loops
Add filters to this data

Mentions Over Time

1,595
Total Mentions
617
Posts with Topic
9
Companies Mentioning
+113.3%
WoW Change

Historical Data

Week Mentions Posts Companies WoW Change
Aug 24, 2026 32 13 9 +113.3%
Aug 17, 2026 15 9 7 -11.8%
Aug 10, 2026 17 11 8 -29.2%
Aug 03, 2026 24 6 4 +14.3%
Jul 27, 2026 21 12 10 -32.3%
Jul 20, 2026 31 15 12 +158.3%
Jul 13, 2026 12 9 6 -50.0%
Jul 06, 2026 24 11 9 +33.3%
Jun 29, 2026 18 10 9 -41.9%
Jun 22, 2026 31 12 9 +158.3%
Jun 15, 2026 12 9 8 +50.0%
Jun 08, 2026 8 4 4 -60.0%
Jun 01, 2026 20 14 10 -23.1%
May 25, 2026 26 13 8 +44.4%
May 18, 2026 18 9 9 -37.9%
May 11, 2026 29 15 9 +16.0%
May 04, 2026 25 11 9 -21.9%
Apr 27, 2026 32 8 5 +28.0%
Apr 20, 2026 25 11 11 +66.7%
Apr 13, 2026 15 10 7 -6.3%
Apr 06, 2026 16 9 7 -40.7%
Mar 30, 2026 27 21 13 +50.0%
Mar 23, 2026 18 9 9 -81.4%
Mar 16, 2026 97 33 22 +223.3%
Mar 09, 2026 30 15 13 +15.4%
Mar 02, 2026 26 13 11 +4.0%
Feb 23, 2026 25 16 12 -28.6%
Feb 16, 2026 35 8 6 -34.0%
Feb 09, 2026 53 21 17 +130.4%
Feb 02, 2026 23 17 14 +27.8%
Jan 26, 2026 18 10 10 -18.2%
Jan 19, 2026 22 10 9 -58.5%
Jan 12, 2026 53 13 11 +20.5%
Jan 05, 2026 44 19 12 +51.7%
Dec 29, 2025 29 5 3 +45.0%
Dec 22, 2025 20 6 5 -44.4%
Dec 15, 2025 36 16 15 -5.3%
Dec 08, 2025 38 17 11 +26.7%
Dec 01, 2025 30 15 11 +200.0%
Nov 24, 2025 10 6 5 -61.5%
Nov 17, 2025 26 14 12 -89.1%
Nov 10, 2025 239 23 13 +895.8%
Nov 03, 2025 24 14 9 -60.7%
Oct 27, 2025 61 16 11 +154.2%
Oct 20, 2025 24 11 9 +140.0%
Oct 13, 2025 10 6 6 -33.3%
Oct 06, 2025 15 6 5 +36.4%
Sep 29, 2025 11 6 5 +22.2%
Sep 22, 2025 9 6 6 -82.4%
Sep 15, 2025 51 8 7 +168.4%
Sep 08, 2025 19 10 9 -9.5%
Sep 01, 2025 21 6 5 +90.9%

Recent Blog Posts

50 posts
Date Company Title Mentions
2026-09-04 Tavily The Shift From Search at Inference to Search in Training 1
2026-09-03 Hugging Face Training a coding model to paint watercolours with TRL and OpenEnv 1
2026-09-02 Baseten Best open-source models for post-training 5
2026-08-31 Fireworks AI Train past the frontier: Training API now generally available 1
2026-08-28 Prime Intellect GLM-5.2 RL weight transfer in 4 seconds using NIXL and ModelExpress 1
2026-08-27 AssemblyAI How to learn machine learning in 2026: an updated roadmap 1
2026-08-27 Hugging Face TAVR: Generate Your Talking Avatar from Video Reference 1
2026-08-27 JetBrains Differential Privacy for Hugging Face Trainers – Without Rewriting Your Trainin… 2
2026-08-26 Fireworks AI Post-training Kimi K3 with Harvey for long-horizon legal work 2
2026-08-25 JetBrains Ideas Worth a Longer Conversation: The JetBrains Research Podcast - The JetBrai… 1
2026-08-25 LaunchDarkly ML Experiment Tracking: What to Track Across Models, Data, and Production 1
2026-08-25 LaunchDarkly Best Practices for Experiment Tracking in MLOps 2
2026-08-25 Anyscale FP8 Reinforcement Learning in SkyRL: Preserving Policy Consistency Across Train… 5
2026-08-25 Hugging Face Granite 4.2 LLMs: How They're Built 12
2026-08-25 Hugging Face Quantization-Aware Healing: a compressed, 4-bit model that outperforms its full… 1
2026-08-25 Lambda AgentFlow: when the agent's workflow learns 2
2026-08-24 Voxel51 AV playbook for robotics: what transfers and what doesn't 1
2026-08-24 Encord Teleoperation vs. Simulation: Where Should Your Robot Training Data Actually Co… 1
2026-08-23 LaunchDarkly ML Experiment Tracking: What to Track Across Models, Data, and Production 2
2026-08-19 Anyscale CVE-2025-62593 and the CISA KEV listing: what Ray users need to know 1
2026-08-19 Encord The Complete Physical AI Data Pipeline: From Data Collection to Deployment 1
2026-08-18 Anyscale Async inference in practice: a video-indexing service on Ray Serve 1
2026-08-18 Anyscale Using Ray Direct Transport for Fast and Easy Weight Syncing in Reinforcement Le… 4
2026-08-18 Redis ReAct agents explained: concepts & practical uses 1
2026-08-17 Crowdstrike Teaching AI to Reason Through Detection Triage 2
2026-08-17 Hugging Face Same Cluster, 33 Points More Utilization: What Changed Was the Order 2
2026-08-17 Lambda A world model for market microstructure 1
2026-08-14 Eden AI GLM-5.3 Benchmark vs GPT-5.6 Sol, Claude Fable 5 & Gemini 3.1 Pro 1
2026-08-14 RunPod The six AI model families and what they're good for 2
2026-08-13 Eden AI Google DeepMind Reshuffle: Hassabis Moves to Chair as Jeff Dean Departs 2
2026-08-13 Hugging Face Sleeper Agents and How to Tame Them 2
2026-08-13 Anyscale Maximizing the Power of NVIDIA GB300 NVL72: NVLink Domain-Aware Placement Group… 1
2026-08-13 Arcee AI Introducing NAC, an Open-Source Harness for Long-Running Agent Work 1
2026-08-12 Patronus AI Getting GLM-5.2 NVFP4 Post-Training off the ground 1
2026-08-12 Hugging Face LFM2.5-VL-3B for Better and Faster Vision Capabilities for the Edge 1
2026-08-11 Baseten Introducing NVIDIA Nemotron 3.5 Lightning 1
2026-08-11 CodeRabbit Teaching NVIDIA Nemotron 3.5 Lightning to route code reviews 1
2026-08-11 Hugging Face Luth-2: Pushing the French Capabilities of SLMs with MOPD 4
2026-08-07 Tavily The Rise of Enterprise Learning Sovereignty 8
2026-08-07 TestMu AI What is AI Model Testing: Methods & Best Practices 1
2026-08-06 Hugging Face FP8 KV-Cache on Intel® Arc™ Pro B70: 2× Capacity with strong Long Context Throu… 1
2026-08-05 AssemblyAI How to build real-time agent assist on streaming speech-to-text 2
2026-08-04 Hugging Face Deploy local agents everywhere with LFM2.5-2.6B 2
2026-08-03 Hugging Face 超越表层对齐:信念是通往深层对齐的新入口 10
2026-07-31 Eden AI Moonshot Distilled Anthropic's Fable: What Model IP Theft and Treasury Sanction… 4
2026-07-31 Eden AI Tiny LLMs That Beat Giant Models: How Efficient 3B-Parameter Models Compete wit… 2
2026-07-31 n8n How LLM Guardrails Keep Production AI Safe 2
2026-07-30 Anyscale Anyscale signs definitive agreement to join Nscale 1
2026-07-30 Arcee AI Teaching an Open Model to Do Science 1
2026-07-29 Braintrust Behavior specs, an open standard for supervising long-horizon agents 1