AgentFlow: when the agent's workflow learns
Blog post from Lambda
AgentFlow, an ICLR 2026 oral paper developed by researchers from Stanford, Texas A&M, UC San Diego, and Lambda, proposes making an AI agent’s orchestration workflow trainable rather than relying on developer-defined prompts and fixed logic for planning, tool use, retrieval, verification, and handoffs. The system separates agents into planner, executor, verifier, and generator modules sharing an evolving memory, while training the planner on-policy through Flow-GRPO, which distributes a final task-success signal across each decision in a multi-turn run to address long-horizon credit assignment. Using a 7B open model, AgentFlow reportedly outperformed larger proprietary systems including GPT-4o across ten search, reasoning, mathematics, and science benchmarks, with gains ranging from 4.1% to 14.9%; in-flow Flow-GRPO training improved results by 17.2%, whereas offline supervised fine-tuning reduced performance by 19.0%. The work argues that reinforcement-learning-based workflow optimization could allow agents to adapt from experience and requires substantial GPU infrastructure for repeated rollouts, policy updates, evaluation, and deployment.
No tracked trend matches for this post yet.
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.