Home / Companies / Observe / Blog / Post Details
Content Deep Dive

Observability Scale: Scaling Ingest to One Petabyte Per Day

Blog post from Observe

Post Details
Company
Date Published
Author
Observe Team
Word Count
2,361
Company Posts That Month
4
Language
English
Hacker News Points
-
Post removed?
No
Summary

Observe offers a data observability platform designed to handle vast amounts of telemetry data, allowing customers to monitor and introspect their systems without needing to sample data, thus supporting up to a petabyte of data ingestion per day for a single tenant. The platform's architecture involves an ingest pipeline that processes data through various stages, including load balancing, data validation, and storage in Snowflake databases, all while ensuring high throughput and cost-efficiency. Initially, Observe struggled with memory and processing challenges as it scaled from handling 30 TB/day to 200 TB/day, but by optimizing their system through changes like converting certain components from Golang to C++ and adopting a parallel datastream insertion strategy, they succeeded in reducing costs and latencies. Further improvements involved addressing Kafka's partition limits and implementing a topic-per-datastream approach for high-volume customers, enabling the platform to manage the ingest of 1 PB/day. Future enhancements are planned, such as utilizing Snowpipe Streaming and refining internal data formats, with the aim of making the system even more efficient and scalable.

Trends Found in this Post
Trend Post Mentions Total Month Mentions Posts Companies MoM
Observability 5 1,257 229 79 +14%
Kubernetes 4 1,735 182 74 +40%
Real-time 2 2,578 595 180 +16%
OpenTelemetry 1 294 39 17 -29%
Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.