November 2023 Summaries
2 posts from WarpStream
Filter
Month:
Year:
Post Summaries
Back to Blog
The new AWS S3 Express One Zone low latency storage class is making waves in the data infrastructure community with its 50% cheaper individual API operations compared to S3 Standard, but it comes at a cost of 8x more per GiB stored. This makes it less suitable as a primary store for big data systems like Kafka and traditional data lakes. However, it opens up an exciting opportunity for modern data infrastructure to tune individual workloads for low latency and higher cost or higher latency and lower cost with the same architecture and code. The new storage class can be used to build completely object storage-based systems with data tiering performed between object storage tiers, potentially reducing costs by an order of magnitude for high volume use cases without touching a line of code. While the initial cost is still high, it's not a non-issue as data can be easily landed into low latency S3 Express buckets and compacted out to S3 Standard buckets asynchronously.
Nov 28, 2023
602 words in the original blog post.
WarpStream is an Apache Kafka protocol compatible data streaming system built on top of object storage, with zero local disks and no inter-zone bandwidth costs. It separates data from metadata, allowing for a massively parallel write engine without synchronization or serialization issues. The system uses a metadata store to track batch sequence IDs, enabling idempotent producer functionality that ensures duplicate batches are dropped before being written to immutable segment files in object storage. This separation of data and metadata also enables "retroactive tombstoning" to identify and drop duplicate batches after they've been written. While implementing idempotency, WarpStream introduced a performance bottleneck due to the need for compaction to merge smaller batches into larger ones, but this was addressed by modifying the file cache interface to support reading batches in a single RPC.
Nov 18, 2023
2,390 words in the original blog post.