Why AI Agents Lose Work When They Crash (and How to Run skills.md Skills Durably)
Blog post from Orkes
AI agents often lose their prompt history, tool results, plans, and approvals when the local process running them crashes, restarts, or is terminated, making simple retry loops insufficient for preserving an entire workflow. The post presents an integration between skills.md, a catalog of more than 200 local and hosted agent skills, and Agentspan, a server-side durable runtime that persists execution graphs, task history, statuses, and approval states independently of client processes. In the example, skills.md CLI calls are wrapped as Agentspan tools to create an incident-reporting agent that parses a CSV file, produces a markdown summary, and generates a PDF, while paid skills require explicit human approval before execution. A simulated out-of-memory crash kills the client container during a run, but the server retains the completed model turn and schedules the pending tool task until a new process on another machine reconnects using the same execution ID and completes the workflow without repeating completed work. Agentspan also provides execution traces showing model turns, tool arguments, timing, token use, and approval pauses, positioning durable state management and auditability as safeguards for production agent workflows.
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.