Ask Miss O11y: Long-Running Requests
Blog post from Honeycomb
When dealing with long-lived streaming RPC workloads, setting service-level objectives (SLOs) can be challenging due to the absence of a clear "success" metric per stream and the potential for streams to last several days. The suggested approach involves instrumenting the workload to provide regular health updates by creating a root span for each stream per minute, known as a "tick," with a span duration of 60 seconds. This setup allows tracking of successful versus failed writes and delays against Kafka offsets, creating metric-like data that can be used to feed SLOs. By aggregating data through a stream ID and using minute-long spans, it becomes possible to monitor each stream's behavior without accumulating excessive spans or waiting for the stream to conclude. This method enables setting SLOs on the number of successful or failed connections per minute and even on individual message success rates, thus ensuring continuous observability and management of the streaming workload.
| Trend | Post Mentions | Total Month Mentions | Posts | Companies | MoM |
|---|---|---|---|---|---|
| Observability | 3 | 696 | 181 | 47 | +0% |
| Real-time | 2 | 984 | 303 | 103 | -12% |
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.