Home / Companies / Redpanda / Blog / Post Details
Content Deep Dive

10X more data, same 4 seconds: single-query scaling in Redpanda SQL on 1 TB

Blog post from Redpanda

Post Details
Company
Date Published
Author
Marcin Grzebieluch
Word Count
2,172
Company Posts That Month
7
Language
English
Hacker News Points
-
Post removed?
No
Summary

Redpanda SQL, a Postgres-compatible analytical query engine, is designed to efficiently handle both live-streaming topics and historical data using bridge queries. Through a benchmark involving NYC Taxi trip records scaled from 100 GB to 1 TB, the study demonstrates that query latency is influenced more by the data scanned rather than the total data retained, with filtering queries taking approximately 4 to 5 seconds regardless of data size. For complex queries requiring full-database scans, performance improves significantly with additional hardware, exemplified by a heavy query's runtime dropping from 298 seconds to 43 seconds when executed on eight nodes. The study suggests that compute cluster sizing should focus on the demands of the heaviest analytical query rather than total data volume, emphasizing that Redpanda SQL's ability to prune data using Iceberg statistics ensures efficient query processing. The findings underscore the importance of distinguishing between scaling for single-query performance and scaling for concurrent query throughput, with further insights to be explored in subsequent discussions.

Trends Found in this Post
Trend Post Mentions Total Month Mentions Posts Companies MoM
Real-time 4 5,522 1,291 230 -4%
Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.