Home / Companies / Select Star / Blog / September 2024

September 2024 Summaries

5 posts from Select Star

Filter
Month: Year:
Post Summaries Back to Blog
Irina Ashurova, Senior Director of Data Development at Pitney Bowes, discusses how Select Star effectively tackled the company's data discovery challenges, particularly in managing large volumes of parcel and mailing event data. The Select Star platform was chosen for its ability to provide simple search functionality, comprehensive data context, visibility of data ownership, and lineage tracking, which surpassed the capabilities of other enterprise solutions previously tested by Pitney Bowes.
Sep 26, 2024 54 words in the original blog post.
Phil Warner, Head of Data at PandaDoc, outlines the company's data transformation efforts as they transition from an outdated Redshift platform to Snowflake, with the goal of overhauling their entire data stack. This shift is driven by reliability issues in the previous system, which undermined user trust and the effective utilization of data. To address these challenges, PandaDoc, a B2B SaaS firm specializing in e-document signing and workflow management, has adopted Select Star as a solution, appreciating its simplicity, clear features, and competitive pricing.
Sep 19, 2024 75 words in the original blog post.
Managing data lakes presents significant challenges, such as data sprawl, inconsistent formats, and lineage tracking issues, which can turn them into data swamps. Apache Iceberg, an innovative open table format, addresses these challenges by offering a robust framework for managing large-scale data sets with improved performance, consistency, and governance. It introduces a new approach to metadata management that enhances query efficiency and data manipulation. The integration with catalog systems supports better governance and access control, ensuring data security and compliance. Organizations like Shopify have benefited from Iceberg's capabilities, achieving reduced data availability latency and enhanced analytics performance. Despite challenges like migrating from legacy systems and balancing real-time ingestion with query performance, Iceberg's features provide a streamlined, efficient, and secure data lake environment. Looking ahead, Iceberg is poised for further advancements, including the integration of table formats and enhanced governance capabilities, making it a significant step forward in data lake technology.
Sep 18, 2024 1,296 words in the original blog post.
Data documentation is crucial for managing growing data complexities within organizations, akin to how libraries evolved from simple systems to structured cataloging like the Dewey Decimal System. As data volumes and complexities increase, effective documentation ensures usability, reliability, and governance, enhancing collaboration, improving data quality, facilitating decision-making, and strengthening governance. The Data Documentation Maturity Curve, divided into Crawl, Walk, and Run stages, outlines a framework for organizations to assess and plan their documentation practices as they grow, recommending tools like Excel for initial stages, dbt Docs for intermediate complexity, and more comprehensive data catalogs like Select Star for advanced stages. These tools help manage metadata, streamline data discovery, provide data lineage, encourage collaboration, and maintain data quality, ensuring that organizations can adapt their documentation practices to support scalable and efficient data management.
Sep 12, 2024 2,170 words in the original blog post.
Effectively managing Snowflake costs is increasingly important as organizations face growing data volumes and evolving usage patterns, with understanding data usage being key to optimizing expenses and improving governance. Strategies discussed by Shinji Kim, CEO of Select Star, at a Snowflake User Group meeting, focus on leveraging data usage context to drive cost optimization, using insights from Select Star to analyze query patterns, joins, and filter conditions derived from Snowflake's account usage and access history views. Key strategies include deprecating unused data and remodeling expensive data pipelines by identifying inactive tables, outdated data, and optimizing resource-intensive queries and data pipelines. Real-world success stories from companies like Pitney Bowes, a fintech firm, and Faire illustrate substantial cost savings achieved through these strategies, such as reducing storage costs, engineering workload, and operational expenses by leveraging usage information and optimizing data models. Implementing these strategies involves regular auditing of data usage, prioritizing high-impact areas, collaborating across teams, and making gradual changes while considering governance and compliance challenges. As data ecosystems evolve, future trends in Snowflake cost optimization may include machine learning for cost prediction, automated recommendations, enhanced cross-platform visibility, and tighter integration with governance frameworks, emphasizing the necessity of continuous monitoring and adjustment.
Sep 03, 2024 992 words in the original blog post.