May 2023 Summaries
22 posts from Starburst
Filter
Month:
Year:
Post Summaries
Back to Blog
Data mesh is a modern data architecture that decentralizes data management, empowering data consumers and supporting various workloads like analytics and AI. It offers benefits such as increased efficiency compared to traditional ETL pipelines, empowerment of individual business domains, enhanced data sharing across domains, and strong data governance and security. However, implementing a data mesh presents challenges, including the need for significant effort, a robust change management strategy, and tailored approaches for individual businesses. It requires focusing on essential business problems, securing buy-in from stakeholders, and deciding the scope of adoption early in the process. Success depends on understanding the unique needs of each organization and effectively communicating the changes and their implications to all relevant teams.
May 30, 2023
930 words in the original blog post.
The Starburst Enterprise 413-e LTS release, announced on May 30, 2023, introduces a range of new features and improvements designed to enhance connectivity, performance, and security for users. This major update, transitioning from a short-term support release to a long-term support version, includes contributions to the open-source Trino project as well as exclusive features for Starburst Enterprise customers. Key enhancements include improved read performance for complex data structures in Parquet, extended support for Delta Lake, and comprehensive support for Apache Iceberg with better statistics generation and query optimization. The release also introduces managed statistics for more connectors, a public preview feature for sharing query tabs, integration with PingIdentity for secure authentication, and enhanced data product security through Apache Ranger. Additionally, fault-tolerant execution is expanded to more connectors, and the release incorporates improvements from the open-source Trino project, such as AWS Lake Formation support and enhanced query editor functionalities.
May 30, 2023
990 words in the original blog post.
Megan Maslanka has joined Starburst as the Chief People Officer, bringing extensive experience in HR and a commitment to building sustainable organizational cultures. Her career began in an international HR role, leading her to pursue an MBA focused on Strategy and I/O Psychology. Megan's journey includes significant roles in tech startups and companies like Buildertrend and DocuSign, where she focused on global expansion and employee experience. She emphasizes the importance of aligning company culture with unique values rather than following trends, advocating for diverse leadership perspectives. Megan, a Nebraska farm native, is driven by hard work and embraces challenges, viewing leadership as a balance of empowerment and accountability. She values character, competence, and ownership, aligning with Starburst's ethos, and is excited about the company's potential in the data lake analytics space. Megan encourages aspiring leaders to embrace their unique perspectives, learn from mistakes, and prioritize organizational growth over individual success.
May 26, 2023
1,710 words in the original blog post.
Trino, originally developed as a fast, interactive query engine to replace Hive, faced limitations when handling batch and ETL workloads, notably encountering a "memory wall" that required either costly scaling or query fragmentation to manage large datasets. To address these challenges, Trino's architecture, which initially relied on massively parallel processing, was re-engineered to incorporate fault-tolerant execution, enabling queries to continue processing despite individual task failures, thus reducing wasted computation and resource use. This new architecture, introduced at Datanova 2023, allows for more efficient resource management, flexible query execution, and the ability to dynamically adjust execution strategies midstream, offering users the ability to run large queries with fewer resources. This advancement is accessible through Starburst Galaxy, where users can explore the fault-tolerant execution mode by creating a free cluster and querying without restarting from scratch in case of task failure.
May 24, 2023
887 words in the original blog post.
Data products offer numerous advantages for both users and businesses by providing a curated and systematic way of interacting with datasets, which enhances collaboration and prioritizes business questions in data analysis. Their adoption is driven by benefits such as ease of use, data reliability, and opportunities for cross-functional sharing, empowering business functions to solve problems efficiently by combining ETL pipelines with business context. Data products democratize data access throughout organizations, enabling any authorized individual to create, save, and share insights, thereby fostering collaboration and reducing misunderstandings. They are designed to be reusable and adaptable, enhancing discoverability and accessibility while supporting data federation, allowing customized data access tailored to specific business needs without duplicating datasets. By promoting committed data ownership and improving user experience, data products enhance organizational processes and problem-solving capabilities, increasing efficiency and autonomy, and facilitating agile workflows. They also bolster security with defined roles and permissions, and their scalability and self-service nature reduce costs and reliance on central IT teams. The integration of Starburst Galaxy's data lake and analytics platform further enhances data product capabilities by offering decentralized data production, a powerful analytical query engine, and robust security features.
May 23, 2023
1,168 words in the original blog post.
Starburst employs a comprehensive approach to securing customer data within its SaaS product, Starburst Galaxy, emphasizing risk mitigation and compliance. The platform's infrastructure is fortified by AWS GuardDuty for continuous monitoring and Cloudflare for protection against DDoS attacks, while access to the user interface is secured with TLS encryption and customer data is never stored within the platform. Payments are processed through Stripe, ensuring credit card information is not collected or stored by Starburst. The company’s practices include rigorous screening of third-party vendors, adherence to GDPR compliance through Standard Contractual Clauses, and a secure development lifecycle that incorporates early consideration of security and privacy, code scanning, and annual penetration tests. Starburst maintains ISO 27001 and SOC2 compliance, striving for continuous improvement in security processes and fostering a culture of feedback to ensure dependable, secure products.
May 23, 2023
544 words in the original blog post.
As AI and data science continue to evolve, the focus has shifted from finding algorithms to effectively scaling applications and making outputs accessible. The rise of data products, which are curated datasets designed to create value for downstream consumers, plays a crucial role in this transition. These data products span both legacy and modern systems and are central to the concept of a data mesh, which decentralizes data ownership across business domains. To efficiently manage data in complex environments, a consumption layer is recommended to decouple end users from data sources, bridging legacy and modern systems while ensuring security through centralized governance. With decentralized access and a data lake as the center of gravity, organizations can economically scale large data volumes and maintain flexibility with open data formats like ORC, Parquet, and Avro. This approach allows for the seamless integration of emerging technologies and enhanced interoperability across the data ecosystem.
May 22, 2023
1,299 words in the original blog post.
Healthcare providers face significant challenges in managing and utilizing vast amounts of data, which can impact patient care, operational efficiency, and financial outcomes. Data from various sources, often fragmented and siloed, complicate efforts to maintain accuracy and continuity in patient records, leading to potential errors and inefficiencies that can be costly both financially and in terms of patient safety. To address these issues, a robust and scalable data analytics platform is crucial, enabling integration, access, analysis, and sharing of data without necessitating complex infrastructure changes. Starburst offers a solution by providing a modern data lake strategy that supports open standards, allowing healthcare organizations to query and analyze data from multiple sources with enhanced speed and security. This approach not only aids in improving patient care and operational effectiveness but also positions healthcare providers to leverage data as a competitive asset. Novant Health, for instance, exemplifies the benefits of using such a platform, achieving faster data access and improved insights into their operations and patient care.
May 18, 2023
1,633 words in the original blog post.
Starburst Galaxy has introduced a general availability connector for Tabular, facilitating the integration of Apache Iceberg and Trino to build modern open data lakes without significant operational overhead. Apache Iceberg, initially developed to address data engineering challenges at Netflix, is a high-performance table storage format that supports SQL operations and solves data lake issues such as schema evolution and data compaction. Starburst and Tabular aim to simplify the management of these open-source technologies, offering tools like access control and data discovery through a managed cloud platform. Tabular provides a secure metastore catalog with role-based access controls, allowing consistent access policies across various compute frameworks. The collaboration between Starburst and Tabular empowers data teams to leverage Iceberg and Trino efficiently, with easy setup through Starburst Galaxy's user interface, and is supported by tutorials and webinars for optimal utilization.
May 17, 2023
783 words in the original blog post.
Apache Hive and Apache Iceberg are two open-source technologies used for managing large datasets, but they differ significantly in their architecture and capabilities. Hive, built on top of Hadoop, allows users to query and analyze big data with a SQL-like interface and is known for its ease of use for non-programmers. However, it faces challenges such as slow file operations, inefficient data manipulation language (DML) operations, costly schema changes, and lack of inherent ACID compliance. Iceberg, designed with modern cloud infrastructure in mind, addresses these limitations by offering efficient updates and deletes, snapshot isolation, and partitioning. It supports full DML on cloud storage, in-place schema changes, and ACID-compliant transactions, making it suitable for different use cases like latency-sensitive data applications, collaborative workflows, root cause analysis, and compliance needs. Migrating from Hive to Iceberg requires careful planning and consideration of specific use cases to optimize data performance effectively.
May 16, 2023
1,262 words in the original blog post.
Data products have become crucial for organizations aiming to leverage their data for competitive advantage, moving beyond being a mere buzzword to a significant advancement in data management. They are defined as reusable data assets tailored for specific uses and delivered according to agreed standards and schedules, ranging from datasets to fraud detection models. Developing successful data products involves focusing on specific, high-impact use cases, assembling multidisciplinary teams with both technical and business expertise, and fostering iterative development for continuous refinement. A robust data product delivery platform is essential for enabling users to discover, understand, and trust data products, supported by comprehensive metadata and data quality measures. Furthermore, establishing automated data governance is critical to maintain data quality, privacy, security, and compliance, which is essential for making data widely accessible without descending into chaos. By adopting these best practices, organizations can ensure that data products remain aligned with business needs and drive analytics success in a data-driven landscape.
May 15, 2023
1,402 words in the original blog post.
Starburst Galaxy has introduced materialized views for catalogs using Great Lakes connectivity, offering a tutorial on using these with Apache Iceberg. Users are encouraged to sign up for Starburst Galaxy, which is free, and connect to a cloud object store to test the features. The tutorial provides a step-by-step guide to creating schemas and an Apache Iceberg table, populating it with data, and creating a materialized view based on a simple query. It explains how the materialized view can be managed and refreshed to ensure it accesses the intended storage table. The process is also applicable to Starburst Enterprise, with further training opportunities available through Starburst Academy.
May 12, 2023
507 words in the original blog post.
This tutorial provides a step-by-step guide on connecting Starburst Galaxy, a managed SQL engine from the creators of Trino, to Tabular, which hosts a sandbox warehouse by default for new organizations. It requires users to have accounts on both platforms, which can be easily set up with free options. The guide details the process of creating credentials in Tabular, setting up a catalog in Starburst Galaxy, and configuring access roles. Once connected, users can add the catalog to a cluster, navigate to the query editor, and explore example tables like the nyc_taxi_yellow table. The tutorial concludes by encouraging users to test queries and learn more about data warehousing in Trino and Apache Iceberg.
May 12, 2023
608 words in the original blog post.
The text discusses the challenges and advancements in anti-money laundering (AML) monitoring within the financial sector, emphasizing the complexity of detecting and preventing financial crimes like money laundering, which accounts for 2-5% of global GDP annually. Financial institutions face difficulties due to the rapid evolution of fraud tactics, the proliferation of financial products, and vast amounts of transactional data that need timely analysis. Starburst offers a solution by integrating its Trino-based platform into existing AML architectures, providing a unified access point to multiple data sources without unnecessary duplication, which enhances data processing, compliance, and security measures. The platform supports real-time analytics and machine learning to identify suspicious activities promptly, aiding in reducing operational costs and risks associated with delayed detection. Additionally, Starburst's innovations, such as Stargate and Warp Speed, enhance performance and allow institutions to meet data sovereignty requirements effectively, while the platform's architecture supports seamless collaboration and adaptability to evolving regulatory demands.
May 11, 2023
2,808 words in the original blog post.
Data products are crucial for effectively managing environmental, social, and governance (ESG) responsibilities, especially as environmental regulations become more stringent. The Greenhouse Gas Protocol defines three scopes of emissions—Scope 1 (direct emissions), Scope 2 (indirect emissions from purchased energy), and Scope 3 (other indirect emissions in the value chain)—requiring organizations to report not only their own emissions but also those of their partners and supply chains. To manage these data demands, a shift from centralized data lakes to a data mesh approach is recommended, where domain experts can identify relevant data for ESG reporting efficiently. This decentralized model, as demonstrated by platforms like Starburst, enhances the searchability, quality, and addressability of data products, enabling organizations to meet ESG requirements effectively. Integrating a data mesh strategy facilitates self-service insights and creates standardized data sets, which are essential for fast and repeatable use across teams, ultimately leading to more informed decision-making and compliance with environmental initiatives.
May 10, 2023
1,080 words in the original blog post.
Trino Fest 2023, a virtual event hosted by Starburst and the Trino community, offers a platform for attendees to engage with the latest innovations and developments in the Trino project. Scheduled to include presentations from industry leaders like Stripe, the festival will feature discussions on new table formats for creating a lakehouse, as well as insights into Trino's integration with tools like Iceberg and Hudi. Keynote speaker Martin Traverso, Trino co-founder and Starburst CTO, will highlight recent accomplishments, while sessions will also showcase various use cases and advancements from the Python community. Participants can register for free, submit speaker proposals, or explore sponsorship opportunities to contribute to the event's success.
May 08, 2023
536 words in the original blog post.
The integration of dbt Cloud with Starburst Galaxy enables the creation of an open data lake architecture, allowing data engineers, analytics engineers, and data analysts to efficiently build, test, and document data pipelines without the need for extensive data migration. This collaboration supports the use of open-source technologies, providing flexibility for businesses to choose between building or buying their data architecture solutions. By leveraging Starburst's capability to federate multiple data sources, users can combine data from various origins, such as AWS COVID-19 data, Snowflake databases, and TPC-H datasets, into a cohesive data lakehouse structure. The process involves reading, cleaning, and optimizing data through different layers—a staging layer for initial data collection, an intermediate structure layer for transformation, and an aggregate layer for final data preparation. The integration simplifies the management of data permissions and enhances accessibility for data consumers, who can view and manipulate aggregated data through role-based access control. The tutorial provided demonstrates setting up a project using dbt Cloud and Starburst Galaxy, showcasing the ease of creating and managing complex data pipelines with these tools.
May 05, 2023
1,533 words in the original blog post.
OS Climate (OS-C) is a collaborative open-source community initiative that aims to enhance global investment in climate change mitigation and resilience by providing a comprehensive data and software platform. Red Hat joined OS-C in 2021 to help integrate climate impact data into financial decision-making and risk management. OS-C addresses the challenge of diverse data sets by developing a federation layer with Trino, allowing access to over a thousand data sources, and providing a self-service infrastructure for teams to tackle climate models and data challenges. They are creating an open-source blueprint to simplify the development of data platforms and promote transparency by allowing contributors to share and enhance data products. OS-C's approach includes managing Environment, Social, and Governance (ESG) data within the constraints of digital sovereignty laws using a federated method. The initiative's ultimate goal is to accelerate the adoption of climate analytics and investment products globally, aligning stakeholder priorities and fostering a community-focused on advancing climate solutions.
May 03, 2023
1,293 words in the original blog post.
Starburst Galaxy has introduced new cluster execution modes to enhance price-performance optimization for data workloads, offering three options: Standard, Fault Tolerant, and Accelerated. Standard mode is the default and is suited for ad-hoc and exploratory analytics, while Fault Tolerant mode, in public preview, allows queries to retry in case of failures, ideal for long-running transformations like ETL processes. Accelerated mode, also in public preview, leverages Starburst Warp Speed's smart indexing and caching to boost query performance by up to seven times, though it is currently limited to four AWS regions. Users can select their preferred execution mode through the Starburst Galaxy user interface, either when creating a new cluster or modifying an existing one.
May 02, 2023
570 words in the original blog post.
Apache Hudi and Apache Iceberg are open-source projects from the Apache Software Foundation that address performance challenges in big data architectures, initially developed to overcome limitations in legacy platforms like Hadoop and Hive. Hudi was created by Uber to reduce data ingestion latency from hours to minutes, while Iceberg, developed by Netflix, was designed to handle ACID transactions and schema evolution, supporting a wide range of file formats and query engines like Apache Spark and Trino. Iceberg enables time travel through its metadata-based approach, capturing snapshots of data states for historical queries and rollbacks. Its scalability and performance make Iceberg a popular choice for data lakehouses, allowing seamless integration with tools like Amazon S3, AWS services, and Snowflake. Starburst Galaxy, leveraging Iceberg, provides a unified platform for managing big data, offering features like federation, near-real-time ingestion, and advanced SQL analytics, which enhance compliance, accessibility, and performance across enterprise data systems.
May 01, 2023
1,421 words in the original blog post.
Cloud object storage has emerged as a cost-effective, scalable alternative to HDFS, transforming how data is managed in cloud computing environments. Unlike HDFS, which stores data in files, object storage manages data in objects, providing architectural distinctions that offer improved storage capabilities and concurrency for distributed data workloads. This technology is particularly beneficial for parallel processing applications, as it allows multiple servers to read data simultaneously, which is enhanced by query engines like Starburst. Cloud object storage, provided by major platforms such as Amazon S3, Microsoft Azure Blob Storage, and Google Cloud Storage, supports dynamic scaling, enabling automatic resource adjustments to meet peak demands and allowing precise resource management. The separation of compute and storage in cloud environments leads to significant cost savings, as users only pay for the resources they utilize. While on-premises installations using HDFS remain relevant, the shift toward object storage reflects a broader industry trend favoring the cloud's flexibility and efficiency.
May 01, 2023
898 words in the original blog post.
A recent study led by Boston Consulting Group, co-sponsored by Red Hat and Starburst, highlights the growing challenges faced by enterprises due to the increasing complexity and costs associated with data architectures. More than 50% of data leaders identify architectural complexity as a significant pain point, exacerbated by the proliferation of data vendors and the exponential growth of data volumes. The report reveals that 50% of stored data is dark, meaning it is not used for actionable insights, while the difficulty in attracting skilled data professionals adds to the complexities. With total data costs projected to grow significantly, organizations are encouraged to adopt new architectural paradigms like data mesh to address issues related to data access and silos. Despite these challenges, the study suggests a shift toward more comprehensive solutions as companies are at a tipping point, urging a reconsideration of current data strategies for better management and insights.
May 01, 2023
1,008 words in the original blog post.