Home / Companies / Starburst / Blog / September 2023

September 2023 Summaries

10 posts from Starburst

Filter
Month: Year:
Post Summaries Back to Blog
The blog post provides a detailed guide on migrating a table from Postgres to a data lake on AWS S3 using dbt Cloud and Starburst Galaxy, leveraging modern table formats such as Apache Iceberg. It outlines the setup process, including configuring catalogs within Starburst Galaxy and integrating dbt Cloud with GitHub, and explains the data transformation steps, focusing on implementing Slowly Changing Dimensions (SCD) Type 2 for handling data changes. The post highlights the use of dbt Cloud for data transformation, including creating snapshot tables and incremental updates, and utilizing Starburst Galaxy's capabilities for efficient querying and metadata tracking. By showcasing a practical example, the guide demonstrates how this migration enables enhanced change data capture and modern data lake functionalities without sacrificing ACID compliance or time-travel capabilities.
Sep 29, 2023 1,528 words in the original blog post.
Starburst Galaxy's tutorial series provides a comprehensive guide for data engineers to build and manage SQL-based data pipelines using modern data lakes. This is part of the Starburst Academy's free course, which emphasizes the simplicity and efficiency of SQL over more complex alternatives like Python UDFs. The tutorial focuses on constructing a modern data lakehouse architecture comprising three layers: Land, Structure, and Consume, using Starburst Galaxy and SQL, with the BlueBikes dataset as a practical example. Participants are guided through downloading the dataset, creating the Land layer to receive raw data, transforming it in the Structure layer, and making it query-ready in the Consume layer. Additionally, the series highlights the integration of Starburst Galaxy with dbt Cloud for automation, enhancing efficiency in data engineering workflows by automating tasks according to a schedule.
Sep 26, 2023 637 words in the original blog post.
Starburst Galaxy, a Software as a Service (SaaS) platform, provides users with the ability to configure clusters tailored to specific workload needs through various execution modes, such as Standard, Fault Tolerant, and Accelerated, optimizing analytics processes. It integrates with over 50 data sources and supports an open data lakehouse architecture using Trino and Iceberg, enhancing data analytics, applications, and AI workloads. Key considerations for configuring clusters include cluster execution modes, sizing, autoscaling, auto suspend, and scheduling, allowing for flexible and efficient management of resources. Warp Speed mode is ideal for interactive analytics with significant data filtering, while Fault Tolerant Execution (FTE) mode suits complex and memory-intensive queries, although it may be slightly slower. Autoscaling adjusts the number of workers based on CPU utilization, and auto suspend conserves resources by shutting down idle clusters. Scheduling features allow clusters to run continuously during specified timeframes, optimizing costs and performance.
Sep 21, 2023 1,394 words in the original blog post.
Trino Summit 2023, organized by Starburst, is set to be a free, virtual educational event taking place on December 13th and 14th, 2023, aimed at sharing the latest developments and use cases within the Trino community. Following the success of Trino Fest, this summit invites speakers and community members to present their work, with Starburst sponsoring the event and encouraging other organizations to do the same to enhance the experience for all attendees. Participants are encouraged to register, propose presentations, and engage with the community through various channels, with the promise of exciting developments and updates to be revealed.
Sep 18, 2023 381 words in the original blog post.
Starburst Enterprise Platform (SEP) 426-e STS is the latest release, following the 423-e LTS version, and is based on Trino 426. This version includes new enterprise features and enhancements, as well as all updates from the recent Trino releases 424, 425, and 426.
Sep 18, 2023 66 words in the original blog post.
Presto, originally developed at Facebook, is an open-source distributed SQL query engine designed for high-performance analytics across various data sources, and it functions without storing data by using connectors for databases, data warehouses, and data lakes. PrestoDB and PrestoSQL (now known as Trino) emerged from the original Presto project, with Trino evolving after its developers formed the Presto Software Foundation for better collaboration and independence. Trino has seen faster development and broader adoption, offering additional connectors and enhanced performance, making it suitable for both ad-hoc queries and batch ETL/ELT workloads, while Presto remains focused on analytics workloads. Both engines utilize ANSI SQL syntax, facilitating integration with existing analytics tools, and are supported by platforms such as Starburst for easy deployment and management. While Presto excels in high-speed analytics, Trino is noted for its fault tolerance and ability to handle large-scale data operations, often offering a more streamlined solution compared to using both Presto and Spark.
Sep 15, 2023 1,300 words in the original blog post.
Data-driven innovation (DDI) is a strategic approach that leverages both known and unknown data, along with known and unknown questions, to drive organizational advancements and decision-making. It stresses the importance of high-quality, well-organized data and a comprehensive data management strategy to support AI and machine learning initiatives effectively. Historically, organizations have relied on centralized data systems, like data warehouses and lakehouses, to address known questions with known data, which often limited exploration and innovation due to the rigidity of traditional data architectures. The text outlines a framework that categorizes data scenarios into four quadrants—known data with known questions, known data with unknown questions, unknown data with known questions, and unknown data with unknown questions—each offering different opportunities and challenges for data exploration and innovation. Emphasizing a shift towards a more flexible and exploratory data ecosystem, the text advocates for utilizing modern data management tools and methodologies that allow for quick experimentation and hypothesis-driven development, which can lead to significant competitive advantages and the creation of new data products. This approach is exemplified by the use of technologies like Starburst that enable rapid data discovery and innovation, facilitating the transition from experimental insights to structured data products that support broader organizational use.
Sep 14, 2023 2,808 words in the original blog post.
Meesho faced challenges in scaling its data infrastructure, prompting a shift from a managed data warehouse to a data lake architecture, utilizing Starburst Enterprise Platform (SEP) and Spark for enhanced performance. Initial deployments of SEP using CloudFormation templates encountered issues with cluster update times and flexibility, leading to a proof-of-concept migration to Kubernetes with SEP Helm charts. This transition aimed to improve fault tolerance and reduce infrastructure costs by leveraging spot instances and employing dynamic scaling strategies. Subsequent design iterations focused on optimizing query execution times, adjusting cluster configurations, and implementing high availability and graceful shutdown processes. The migration to Kubernetes enabled more aggressive scaling, reduced software costs, and maintained low execution times, ultimately achieving a 30% cost reduction and faster deployment times compared to the previous AWS CloudFormation-based setup.
Sep 12, 2023 1,897 words in the original blog post.
Starburst Galaxy has introduced support for Python DataFrames through the PyStarburst and Ibis libraries, allowing data engineers to perform complex data transformations using Python while leveraging the performance of the Trino SQL query engine. This integration aims to streamline development practices by enabling a unified platform for both analytical and transformation workloads, eliminating the need for separate engines and reducing associated costs and complexity. PyStarburst allows for easy migration of existing PySpark and Snowpark workloads, facilitating the use of Python DataFrames within Starburst, while Ibis provides a uniform Python API that decouples execution from the DataFrame API, enabling scalable data processing across various data sources. This development is part of Starburst's broader effort to simplify data engineering processes and support open-source initiatives, in collaboration with partners like Voltron Data.
Sep 07, 2023 1,026 words in the original blog post.
Starburst Enterprise 423-e LTS introduces a suite of enhancements aimed at improving connectivity, performance, security, and storage capabilities for its users. The release bolsters data management with new features for Apache Iceberg, Delta Lake, and Hive, including improved version control, data lineage tracking, and query optimization. Performance gains are achieved through advanced indexing, column organization, and support for new instance types, while security is enhanced with SCIM integration for identity management and updates to AWS Lake Formation. Connectivity sees the introduction of a parallel Snowflake connector, significantly boosting data integration speed, and managed statistics are expanded to more connectors to optimize query performance. Storage advancements include support for Ceph, Dell ECS, and ObjectScale, offering users greater flexibility and the ability to manage multi-cloud data analytics efficiently. These improvements are complemented by contributions to the open-source Trino project, ensuring a comprehensive and robust data platform.
Sep 07, 2023 1,374 words in the original blog post.