February 2023 Summaries
12 posts from Starburst
Filter
Month:
Year:
Post Summaries
Back to Blog
In a discussion on Data Mesh TV, Uday Hegde, CEO of USEReady, elaborates on the company's Migrate, Optimize, Modernize (MOM) framework, designed to assist organizations, particularly in heavily regulated industries, in their data transformation efforts. The MOM framework involves three phases: 'Migrate' focuses on moving data efficiently while minimizing business disruption, 'Optimize' aims to create cost-effective, user-friendly data products, and 'Modernize' ensures that data capabilities are fully utilized and trusted by consumers. USEReady employs tools like MigratorIQ, OptimIQ, and DecisionIQ to guide these processes, with an emphasis on cloud automation, self-service capabilities, and cultural alignment. The conversation highlights the importance of agility, quick iteration, and leveraging AI to address data quality challenges, advocating for community collaboration and knowledge sharing in the data mesh community.
Feb 28, 2023
964 words in the original blog post.
Starburst Enterprise's 407-e LTS release introduces significant enhancements such as a managed statistics feature that optimizes query performance for databases like Oracle, PostgreSQL, and Teradata by collecting table and column statistics. The release also expands fault-tolerant execution to include write operations with MongoDB and BigQuery, and it now supports HDFS storage as an external spooling option. Starburst Warp Speed, a feature that enhances query performance on data lakehouses through indexing and caching, is now generally available with the Starburst Enterprise Elite license, including new REST endpoints and a public preview of index and cache resiliency features. Additionally, the Starburst Enterprise REST API has been updated to facilitate programmatic management of built-in access control, and support for AWS Lake Formation access control has been extended to offer more granular options, such as Glue user impersonation and AWS role selection per catalog. These updates are part of ongoing efforts to improve the platform, with detailed release notes available for users seeking comprehensive information.
Feb 28, 2023
376 words in the original blog post.
At the Datanova event, Justin Borgman's keynote addressed common misconceptions about big data over the past decade, offering truths to correct these narratives, which resonated with Adrian Estala, Starburst's VP and Field Chief Data Officer. Estala reflected on how these misconceptions have affected Chief Data Officers (CDOs) and their business commitments, advocating for a positive shift in strategy rather than a complete reset. CDOs are encouraged to focus on providing immediate access to data assets and enabling self-service capabilities, thereby saving on migration costs and reinvesting in analytics projects. The emphasis is on integrating data from existing and new sources, simplifying processes, and empowering users with self-service tools to improve data-driven initiatives, especially AI and ML, which often suffer from delays in accessing trusted data. Estala stresses the importance of a modern data ecosystem over a technology stack, highlighting the need for rapid integration and iteration to unlock tangible business value.
Feb 23, 2023
595 words in the original blog post.
Shopify significantly improved its data processing efficiency by migrating from Hive and JSON table formats to Apache Iceberg and Parquet, using Trino as the compute engine. This transition, prompted by data silos and interoperability challenges, led to execution time reductions from hours to mere minutes, greatly enhancing the productivity of data analysts and scientists. The migration involved rewriting vast amounts of data and overcoming technical challenges with the help of the Trino community, showcasing the benefits of modern table formats and open-source collaboration. The shift resulted in execution speeds that were 1000 times faster, underscoring the importance of adopting efficient data storage and processing frameworks to drive business insights and performance.
Feb 21, 2023
1,022 words in the original blog post.
Comcast's journey into implementing a data mesh with Starburst marks a significant evolution in its data strategy, aiming to harmonize its data ecosystem while maintaining business unit independence. Initially pioneering data virtualization, Comcast now utilizes Starburst to build a data mesh that integrates storage, computation, and transformation, emphasizing the cultural shift required to effectively support this system. This approach facilitates interoperability between different data storage solutions without necessitating costly migrations, allowing Comcast to leverage existing infrastructure while adopting new technologies. The transition to data mesh, supported by comprehensive user education, simplifies data governance and enhances privacy and security compliance, aligning with regulations like GDPR and the California Consumer Privacy Act. By promoting a unified data governance model and improving performance, Comcast demonstrates the potential for data mesh to become a fundamental component of enterprise data strategies, supporting a seamless user experience and enabling quicker technological adaptations.
Feb 14, 2023
1,335 words in the original blog post.
In his Datanova talk, Benn Stancil examines potential failures in data engineering by drawing parallels with the film "World War Z," emphasizing the need to anticipate scenarios where popular tools like Snowflake, Fivetran, and DBT might not succeed as expected. Stancil identifies several challenges, including the monotony of data engineering tasks, the high costs associated with data engineers and tools, and the risk of roles being replaced by automation or decentralized data management trends such as data mesh. He also discusses the possibility of AI automating many data engineering functions, potentially transforming the role into one more focused on infrastructure maintenance. Stancil advocates for a shift from technology-centric to problem-centric approaches, urging the data engineering community to focus on creative, empathetic problem-solving by understanding real-world issues faced by users rather than solely relying on technological solutions. This approach would involve listening to the broader community's needs to successfully navigate the evolving landscape of data engineering in the age of AI.
Feb 11, 2023
1,419 words in the original blog post.
The 2023 Data Rebel Awards, presented by Starburst at their annual Datanova event, recognized outstanding contributions in data, analytics, and AI across several categories, celebrating individuals and organizations that embody the spirit of innovation and leadership in the field. Honorees included Chandrasekhar Vemuri for his groundbreaking work in data access at Société Générale, Richard Jarvis for improving healthcare data efficiency at EMIS Health during the COVID-19 pandemic, and Caroline Chung for advancing data-driven cancer care at MD Anderson Cancer Center. The awards also acknowledged achievements in data virtualization, analytics strategies, data architecture, and engineering, with winners like Venkata Bitra, Ludovic Staehli, and Benjamin Jeter making significant impacts in their respective fields. Additionally, partner awards highlighted the contributions of companies like Accenture, Slalom, and Google, recognizing their roles in driving technological advancements and fostering partnerships within the data ecosystem.
Feb 08, 2023
1,742 words in the original blog post.
Vendor-provided benchmarks, often used to demonstrate the performance of database systems, can be misleading as they do not accurately reflect real-world production environments. These benchmarks, such as TPC-H and TPC-DS, assume ideal conditions where data is already optimized and centrally located, ignoring the complexities and costs associated with data preparation and transfer. As data strategies evolve towards distributed architectures like data lakes and data meshes, traditional benchmark metrics fail to capture the true cost/performance balance experienced by users. Real-world workloads are diverse, involving concurrent queries of varying sizes that strain system resources and impact performance. Vendors may skew benchmark results through practices like excessive caching, which are not feasible in large-scale, real-world scenarios. Organizations are encouraged to conduct their own performance tests to obtain vendor-neutral results that better reflect their specific production needs. Starburst, with its data source agnostic capabilities, advocates for this approach and supports the use of open data formats and diverse architecture choices to optimize performance and cost-effectiveness.
Feb 07, 2023
767 words in the original blog post.
Businesses striving to become data-driven often mistakenly believe that bridging the skills gap requires hiring more engineering talent or upskilling their current workforce, but this approach is costly and ineffective. Many companies have realized they cannot compete with high salary offers for top engineering talent, leading them to focus on internal upskilling, which is challenging and resource-intensive. Instead, technology can serve as a more efficient bridge between analysts and data, enabling powerful AI and analytics capabilities through user-friendly interfaces that automate much of the necessary work. This approach allows IT teams to focus on higher-value tasks, while ensuring accurate data is fed into models, ultimately making businesses more data-driven without the need for extensive new hires. Starburst positions itself as a solution to this challenge by providing a platform that supports self-service analytics, allowing companies to connect analysts and engineers more cost-effectively and efficiently.
Feb 06, 2023
661 words in the original blog post.
Many companies eager to become AI-driven enterprises often make the mistake of prioritizing hiring top-tier data scientists and investing in the latest AI technology without first establishing a solid data foundation. This misstep can lead to inefficiencies, as data scientists end up performing mundane data management tasks, and AI models falter without robust data to support them. The sudden shift during the pandemic highlighted the inadequacy of relying solely on historical data and underscored the necessity for real-time analytics infrastructure. To truly leverage AI's potential for boosting revenue, reducing costs, and enhancing product experiences, businesses must focus on improving data management and access to high-quality data. Starburst offers a solution by providing comprehensive access to data across an organization, enabling data scientists to focus on developing effective AI models. Investing in a strong data infrastructure now can profoundly impact a company’s long-term success in the evolving AI landscape.
Feb 03, 2023
661 words in the original blog post.
The notion of the "modern" data stack as a revolutionary advancement is challenged by the argument that it is merely a repackaging of traditional data strategies, now optimized for cloud environments without significant architectural change. While the cloud-based approach offers scalability, flexibility, and cost-effectiveness, it perpetuates the centralization of data akin to previous systems. The core principles remain unchanged, and the promise of modernization through migration is questioned, suggesting that true transformation requires a holistic evolution of data strategies beyond mere technological updates. Starburst advocates for a future where data systems prioritize speed, scalability, simplicity, and SQL, aiming to reduce complexity and latency while being vendor-agnostic. This approach underscores the need for businesses to efficiently connect and analyze diverse data sources to drive meaningful insights and decisions.
Feb 02, 2023
1,846 words in the original blog post.
The concept of a "single source of truth," where all of a company’s data is centralized in one location, has been marketed by technology vendors such as Snowflake, SAP, and Oracle as an ideal solution for businesses to fully understand their operations. However, this approach is criticized for being impractical and costly, as it often results in outdated data by the time it is centralized and fails to accommodate the immediate insights needed for modern businesses. Instead, Starburst offers a platform that allows data to remain in its original location, providing a secure yet flexible means of accessing and querying data without the need for centralization. This approach facilitates rapid data access, compliance with data privacy laws, and enables engineers and analysts to focus on more critical tasks, challenging the traditional centralized data architecture promoted by legacy vendors.
Feb 01, 2023
574 words in the original blog post.