Home / Trends / Reinforcement learning

Reinforcement learning

Reinforcement and preference learning for AI models

Explore historical trend data

Sign up for free to navigate to earlier time periods.

Tracked keywords: reinforcement learning, reward-based learning, rlhf, agent learning via feedback, preference learning, fine-tuning loops
Add filters to this data

Mentions Over Time

788
Total Mentions
334
Posts with Topic
6
Companies Mentioning
-36.8%
WoW Change

Historical Data

Week Mentions Posts Companies WoW Change
Jul 13, 2026 12 9 6 -36.8%
Jul 06, 2026 19 9 8 +46.2%
Jun 29, 2026 13 7 6 -48.0%
Jun 22, 2026 25 8 6 +150.0%
Jun 15, 2026 10 7 6 +11.1%
Jun 08, 2026 9 5 5 -30.8%
Jun 01, 2026 13 9 8 -40.9%
May 25, 2026 22 10 6 +69.2%
May 18, 2026 13 6 6 -45.8%
May 11, 2026 24 12 6 +26.3%
May 04, 2026 19 7 6 -40.6%
Apr 27, 2026 32 8 5 +33.3%
Apr 20, 2026 24 10 10 +166.7%
Apr 13, 2026 9 6 4 -10.0%
Apr 06, 2026 10 5 4 -54.5%
Mar 30, 2026 22 16 10 +57.1%
Mar 23, 2026 14 6 6 -74.5%
Mar 16, 2026 55 21 14 +129.2%
Mar 09, 2026 24 12 9 +14.3%
Mar 02, 2026 21 8 6 -12.5%
Feb 23, 2026 24 15 11 -31.4%
Feb 16, 2026 35 8 6 -16.7%
Feb 09, 2026 42 16 14 +100.0%
Feb 02, 2026 21 15 12 0%
Jan 26, 2026 21 10 10 0%
Jan 19, 2026 21 9 8 -56.3%
Jan 12, 2026 48 11 9 +14.3%
Jan 05, 2026 42 18 11 +50.0%
Dec 29, 2025 28 4 2 +47.4%
Dec 22, 2025 19 5 4 -42.4%
Dec 15, 2025 33 13 12 -10.8%
Dec 08, 2025 37 16 10 +37.0%
Dec 01, 2025 27 13 9 --

Recent Blog Posts

50 posts
Date Company Title Mentions
2026-07-16 Hugging Face One Adapter, Both Modalities: Field Notes from Building and Serving a Multimoda… 1
2026-07-15 Hugging Face Welcome Inkling by Thinking Machines 3
2026-07-15 Hugging Face Introducing Real World VoiceEQ: Measuring the human quality of voice AI 1
2026-07-14 NeuralTrust AI Token Optimization: Complete Guide to Reducing LLM Costs 1
2026-07-14 NeuralTrust AI Token Optimization: Complete Guide to Reducing LLM Costs 1
2026-07-14 Hume Introducing Real World VoiceEQ: Measuring the Human Quality of Voice AI 1
2026-07-13 Eden AI GPT-5.6 Sol Ultra and the Cycle Double Cover Conjecture: AI Reaches a Mathemati… 1
2026-07-13 Arize How do you make an LLM, anyway? Microsoft just published a textbook. 2
2026-07-13 Parallel Web Systems Introducing Parallel Search Turbo 1
2026-07-11 Braintrust Best LLM fine-tuning platforms in 2026 1
2026-07-10 Cohere Hardware-aware dynamic speculative decoding 1
2026-07-09 Encord AI Data Curation for LLM and Multimodal Teams: A Practical Framework 1
2026-07-08 Cursor Introducing Grok 4.5 2
2026-07-07 Encord What Is AI Data Labeling? Definition, Types and the Process 8
2026-07-07 Hugging Face From Hugging Face to Amazon SageMaker Studio in one click 2
2026-07-07 JetBrains The Benchmark Meaning Gap - The JetBrains Blog 1
2026-07-06 PromptLayer Top 5 Chinese LLMs: The Models Powering China’s AI Surge in 2024–25 1
2026-07-06 Daytona GPU Sandboxes 2
2026-07-04 Redis Multi-step AI agents: what they are & how they work 1
2026-07-03 PromptLayer How to Build Your Own Deep Research 3
2026-07-01 Deepinfra MiMo-V2.5 Model Documentation and Integration Guide 2
2026-07-01 Hume Emotional Intelligence Is a Training-Time Property, Not a Prompt 3
2026-06-30 Hugging Face Why Specialization Is Inevitable 1
2026-06-30 Datadog Datadog acquires Adaptive ML 1
2026-06-30 Together AI Together AI at ICML 2026: frontier research across the full stack 2
2026-06-30 TestMu AI What is One-Shot Prompting: A Complete Guide 1
2026-06-29 TestMu AI What is Zero-Shot Prompting: A Complete Guide 3
2026-06-28 Resemble AI Proteus: Automated Adversarial Robustness Testing for Audio Deepfake Detectors 1
2026-06-26 Fireworks AI Cursor Composer 2 + Fireworks AI 10
2026-06-26 SSOJet 10 AI QA Agents That Test Code Before You Ship 1
2026-06-26 SSOJet 7 AI Test Generation Tools for Developers 3
2026-06-25 Baseten Live draft model training for speculative decoding 1
2026-06-25 NeuralTrust Chain-of-Thought Hijacking: How Longer Reasoning Breaks AI Safety 3
2026-06-24 AssemblyAI Wrong drug name in, wrong SOAP note out: error propagation in clinical AI pipel… 1
2026-06-24 Fireworks AI Frontier-lab Training Infrastructure, Available Now as a Managed Service for GL… 4
2026-06-24 Hugging Face Interhuman’s Goblin: “Yeah, Friday at Five” 1
2026-06-24 LabelBox Introducing Recursion: The RL platform for enterprise specialist agents 3
2026-06-24 Gradium Launching Gradium Translate: the best accuracy-latency tradeoff against gemini-… 1
2026-06-22 Hugging Face V-Zero 2
2026-06-18 Bright Data Embodied AI in 2026: Everything You Need to Know 4
2026-06-18 Bright Data VLAs and World Models Need Web-Scale Data. Just Not the Same Data 1
2026-06-18 Vast.ai The Future of AI Inference in 2026: Key Trends Shaping AI Infrastructure 1
2026-06-17 JetBrains Step Rejection Fine-Tuning: Squeezing More Signal from Noisy Agent Trajectories… 1
2026-06-17 Roboflow Physical AI: How AI Systems Interact with the Physical World 1
2026-06-16 Anyscale Data Processing is Becoming a GPU Workload 1
2026-06-15 Cloudflare Growing the Cloudflare AI team with talent from Ensemble AI 1
2026-06-15 Hugging Face How We Built OpenMythos: A Cybersecurity LLM Trained from Scratch 1
2026-06-11 Anyscale How Adyen trains a Transaction Foundation Model (TFM) on 51 trillion tokens and… 2
2026-06-10 NeuralTrust Unmasking the Machine: A Technical Deep Dive into AI Identity Disclosure 1
2026-06-09 Qovery Coding Agents Write the Code. Who Verifies It Works? We Built the Answer. 1