Home / Companies / Starburst / Blog / January 2026

January 2026 Summaries

8 posts from Starburst

Filter
Month: Year:
Post Summaries Back to Blog
A semantic layer serves as a business-friendly intermediary that translates complex data models into understandable metrics and dimensions for business users, ensuring consistent and governed access to data across various tools and applications. It bridges the gap between raw data storage and consumption by connecting data lakes, warehouses, and streaming systems to BI tools and AI systems, which require structured business context. The need for a semantic layer has grown as organizations face the challenge of inconsistent data metrics across multiple BI tools, leading to a lack of trust in data and increased manual aggregation efforts. The adoption of semantic layers is further driven by AI's requirement for structured data and the need for operational activation beyond traditional analytics. However, implementing a semantic layer can be technically challenging due to fragmented standards, potential performance bottlenecks, governance complexities, and the need for robust change management. Successful implementation involves a pragmatic strategy, starting with high-impact use cases, designing for performance and governance, and treating semantic models as data products. By focusing on open standards and integrating advanced optimization features, organizations can build scalable semantic layers that meet current needs and adapt to future requirements, as demonstrated by case studies in various industries.
Jan 30, 2026 1,569 words in the original blog post.
Daniel Abadi's study explores the significant performance bottleneck caused by using JDBC connections in Trino when extracting large amounts of data from traditional database systems like Oracle, MySQL, and PostgreSQL. Through experiments using the TPC-H benchmark, Abadi demonstrates the inefficiencies of JDBC, highlighting a stark contrast in data access speed when comparing JDBC-based extraction to non-JDBC methods, such as those involving HDFS. The experiments reveal that parallel JDBC connections significantly enhance performance, as shown by the use of Starburst's advanced connectors for Oracle, which achieve substantial speedups by leveraging data partitioning. While Starburst provides upgraded connectors for parallel data extraction, alternative strategies, like logical partitioning in Trino, can also mitigate the JDBC bottleneck. Despite these challenges, Abadi underscores Trino's capability to access and process federated data across multiple systems, advocating for strategic actions to optimize performance by addressing the JDBC bottleneck, thereby fully utilizing Trino's potential in accessing data from diverse sources.
Jan 27, 2026 2,657 words in the original blog post.
Data products and data lakehouses are two complementary components that together address the complexities of modern data management, particularly in large organizations with distributed data systems. Data products apply a customer-centric approach to data, emphasizing discoverability, governance, and ease of use, while data lakehouses offer a technical foundation that combines the flexibility of data lakes with the performance of data warehouses. This combination allows for federated data access without requiring all data to be centralized, thus preserving data locality and enabling efficient data consumption across various domains. Key technical enablers like Apache Iceberg and Trino facilitate schema evolution, query federation, and integration between batch and streaming data, ensuring that data products remain adaptable and accessible. These capabilities are particularly beneficial for AI and machine learning workloads, which require context-rich data and robust governance to meet regulatory standards. By enabling distributed but coordinated governance, data products and lakehouses allow organizations to manage data efficiently across multiple systems and regions, reducing data movement costs and enhancing data freshness. This architecture is increasingly adopted across industries like financial services, healthcare, and government for applications ranging from risk management to secure data sharing, offering a scalable and unified approach to data strategy that supports both current analytics and emerging AI needs.
Jan 23, 2026 1,249 words in the original blog post.
Cloud data warehouses (CDWs) like Snowflake, BigQuery, and Redshift have been central to modern analytics due to their ability to handle structured BI tasks efficiently. However, when data exceeds traditional clean table formats and shifts towards raw, semi-structured, or real-time applications, these warehouses can become costly and less effective. The rise of open lakehouse architectures, combining object storage with open table formats like Apache Iceberg, offers a flexible alternative by allowing diverse compute engines to operate on the same data. This architecture supports various workloads, including interactive SQL, batch processing, and machine learning, often at a lower cost due to its scalable object storage model. As organizations face rising costs, concurrency issues, and workload mismatches in their existing CDWs, many are considering a hybrid approach that retains CDWs for polished BI tasks while transitioning other workloads to a lakehouse setup. This strategic shift helps optimize data performance and cost-efficiency without an abrupt overhaul of existing systems.
Jan 20, 2026 1,637 words in the original blog post.
Apache Iceberg is gaining popularity as an open-source table format that offers high performance, schema evolution capabilities, and full CRUD support, making it suitable for various workflows, including analytics and AI. While transitioning to Apache Iceberg can seem daunting due to concerns about data migration, maintenance, governance, and training, it is not an all-or-nothing proposition. Organizations can gradually migrate critical workloads by leveraging modern distributed query engines like Trino, which support Iceberg natively, allowing for a flexible approach with selective centralization. Apache Iceberg's metadata utilization supports features like time travel, rollback, and snapshots, making it ideal for modern data platforms with growing, changeable needs. To facilitate adoption, companies should engage stakeholders, address organizational constraints, and establish a maintenance strategy to ensure long-term performance. Starburst Galaxy offers a managed platform to ease the migration process, providing tools that simplify maintenance, automate tasks, and integrate with leading data platforms, thereby minimizing the complexity of adopting Apache Iceberg.
Jan 15, 2026 1,925 words in the original blog post.
A data lakehouse, particularly one built on the Iceberg architecture, offers a solution for organizations managing complex data estates that have outgrown traditional data warehouses and lakes, especially when faced with escalating storage costs, fragmented governance, and vendor lock-in. Combining the data warehouse's structured query capabilities with the data lake's flexibility, a lakehouse facilitates unified governance, cost-effective storage, and seamless access to both structured and unstructured data. This architecture supports advanced analytics and AI initiatives by providing a single access point for BI and data science workloads, reducing the need for multiple data engineering efforts and eliminating inefficiencies caused by duplicated pipelines and approvals. Organizations can future-proof their data strategies by leveraging open standards like Apache Iceberg, allowing them to integrate with diverse compute engines such as Trino and Spark, while maintaining strong governance through consistent role-based access controls and audit trails. This approach not only addresses the challenges of high data warehousing costs and governance fragmentation but also accelerates data-driven innovation by enabling faster and more flexible data processing and analysis.
Jan 13, 2026 1,243 words in the original blog post.
Federated search engines address the challenges of distributed data environments by enabling SQL queries across heterogeneous data sources without the need to move data, thus alleviating the operational friction caused by data silos. These silos emerge when different business units use independent technology stacks, often exacerbated by mergers, acquisitions, and regulatory frameworks like GDPR. Traditional centralization efforts face issues such as high storage costs and data staleness, whereas federated queries maintain data in its original location, reducing infrastructure demands and improving compliance by adhering to data sovereignty regulations. Modern federated engines use massively parallel processing to enhance query performance across distributed systems, leveraging native optimizations of each source, and thereby overcoming the limitations of early federation systems. This approach also streamlines the onboarding of new data sources, shortens time-to-insight, and supports multi-cloud and hybrid architectures by querying across cloud boundaries without vendor lock-in. Nonetheless, federated queries require careful technical implementation, including connector configuration, query optimization, and security architecture, while ensuring consistent data governance across distributed systems.
Jan 08, 2026 1,648 words in the original blog post.
Starburst Galaxy is a fully-managed Trino service built on a multi-cloud lakehouse platform that enhances data architecture by simplifying deployment, scaling, and security, allowing teams to focus on analytics rather than operational tasks. It operates through three interconnected planes—control, service, and data—across AWS, Google Cloud Platform, and Microsoft Azure, using Kubernetes clusters for low-latency access and workload isolation. The control plane manages identity, policy, and infrastructure coordination, while the service plane handles query requests and the data plane ensures elastic execution of SQL queries. Starburst Galaxy also emphasizes robust data governance with unified policy models, data masking, and lineage auditing to ensure secure and efficient data handling. By leveraging Apache Iceberg and providing integration with tools like dbt, it offers a familiar SQL experience and supports interoperability, avoiding vendor lock-in. Starburst Galaxy is designed to cater to the needs of data engineers and analysts seeking a streamlined solution with advanced features, and it continuously evolves to incorporate AI and other innovations for enhanced productivity.
Jan 08, 2026 1,613 words in the original blog post.