Home / Companies / Acceldata / Blog / March 2022

March 2022 Summaries

11 posts from Acceldata

Filter
Month: Year:
Post Summaries Back to Blog
LinkedIn operates one of the largest analytics platforms in the world, with approximately 20,000 Hadoop nodes storing one exabyte of data. The company's largest data cluster comprises 10,000 nodes storing 500 PB of data including one billion objects (directories, files and blocks), all using the Hadoop Distributed File System (HDFS). LinkedIn has developed various tools to optimize its storage and compute performance, such as Azkaban for workflow job scheduling, Robin for load balancing, and DynoYARN for performance forecasting. The company is currently in the process of migrating its on-premises Hadoop clusters to the Azure public cloud owned by parent Microsoft.
Mar 31, 2022 2,308 words in the original blog post.
Optimizing enterprise-scale data infrastructure and operation costs is a significant challenge, with 82% of organizations running cloud infrastructure workloads incurring unnecessary costs. Data observability helps data teams optimize their data at scale and achieve incredible ROI on their data investment. By providing a unified view of all data systems across the entire data lifecycle, data observability can help enterprises reduce downtime, improve data quality, and save on software and hardware expenses. Acceldata's suite of data observability solutions has helped companies like PhonePe, True Digital, and PubMatic to save several million dollars a year as they scaled their data infrastructure.
Mar 31, 2022 1,983 words in the original blog post.
In a recent Modern CTO podcast, Acceldata's CTO and Co-founder, Ashwin Rajeev, discussed the company's journey from its inception in 2018 with just a handful of employees to now being a global data leader with over 100 workers. The key challenge for Ashwin and his team was managing hypergrowth across operations, sales, and engineering while addressing each department's specific needs. A crucial aspect was hiring talented individuals who could find meaning in their work by making an impact. Acceldata encourages its employees to contribute immediately and fosters a culture that emphasizes customer-centric problem-solving.
Mar 30, 2022 221 words in the original blog post.
Companies are racing to become fully-fledged 21st-century data-driven businesses, with some leading the pack and others lagging behind. However, many companies focus too much on building powerful data engines but neglect investing in data observability tools that provide visibility into their data operations. This lack of insight can lead to massive blind spots, causing companies to fall behind competitors. Businesses need multi-dimensional data observability solutions to ensure data flows smoothly and is error-free wherever it travels, stored, or processed. By eliminating data blind spots, companies can improve data speeds, reduce unplanned outages, and save costs on bandwidth, storage, software licensing, and system processing.
Mar 29, 2022 1,526 words in the original blog post.
Researchers found that improving data quality and usability by 10% could increase return on equity (ROE) for Fortune 1000 companies by 16%. To achieve this, enterprise data teams need a data observability solution with advanced AI/ML capabilities to automatically detect data and schema drift, anomalies, as well as lineage. Data observability offers full traceability of how data transforms across the entire data lifecycle, helping predict, prevent, and resolve unexpected data downtime or integrity problems. Acceldata Torch is a multi-dimensional data observability solution that provides a single unified view of the entire data pipeline across different technologies throughout the entire data lifecycle. It can help ensure data reliability even after the data transforms multiple times across several different technologies and automatically identify anomalies, root causes, and classify large sets of uncategorized data. AI and ML capabilities are crucial for enterprises to improve data quality at scale as manual interventions alone aren't sufficient.
Mar 22, 2022 1,154 words in the original blog post.
In a recent episode of The Data Engineering Podcast, Tristan Spaulding, Head of Product at Acceldata, discussed the concept of data observability and its importance in modern data environments. He emphasized that to operationalize data effectively, a multidimensional approach is necessary, focusing on understanding computational and logical elements powering analytical capabilities. Data teams should have accurate information about their data at all times, with continuous observation, operation, and optimization capabilities enhanced by automation and machine learning. Reliable data pipelines supported by ML models are crucial for navigating the modern data stack, cloud shift, and legacy data platforms.
Mar 18, 2022 222 words in the original blog post.
Cloud adoption is rising among enterprises due to benefits such as lower costs, scalability, agility, and resilient operations. However, migrating data and workloads to a cloud environment requires a thoughtful plan that aligns with an organization's data strategy. Data observability can help create effective SLAs/SLOs for external service providers and optimize data operations. It provides performance, usage, and cost baselines that inform all cloud migration decisions. Understanding the different types of clouds (public, private, hybrid) is important when choosing which one(s) is right for a business. Cloud migration strategies include rehosting, replatforming, repurchasing, refactoring, retiring, and retaining. Acceldata's Data Observability Platform can help optimize data spend, improve operational intelligence, and ensure data reliability during cloud migrations.
Mar 17, 2022 1,541 words in the original blog post.
Acceldata Pulse On-Prem 2.0 has been released with new features including Hydra for secure access to cluster nodes and Alerts enhancements such as new notification channels, predefined alerts, and a new Kafka alert category. Other improvements include Kafka capacity planning, Impala resource pool properties, Pulse CLI additions, and new runbooks for HDFS balance and DNS lookup checks. Additionally, error logs retention has been extended, and the Nodes page now supports sorting by host in its heat map chart.
Mar 15, 2022 1,475 words in the original blog post.
The text discusses the growing importance of data observability as a distinct product category for enterprises focusing on data optimization and utilization. It highlights the evaluation criteria for identifying essential features in a data observability platform, including data sources and collectors, monitoring and measuring data environments, analyzing and optimizing data pipelines and infrastructures, operating and managing data operations, and assessing data architecture. The text also covers non-technical aspects such as market viability and pricing models for vendors in this space.
Mar 14, 2022 3,436 words in the original blog post.
JP Morgan Chase has developed a comprehensive data structure based on the concept of "data products" to manage its vast amounts of data. The bank uses Amazon's cloud services, including AWS Glue Data Catalog and Lake Formation, to create a decentralized data mesh architecture that allows for secure sharing of data across the organization while maintaining control over it. This approach helps JP Morgan Chase achieve cost savings, business value, and data reuse.
Mar 08, 2022 859 words in the original blog post.
The text discusses the importance of reliable data in decision-making processes, emphasizing that good data quality is not enough. It suggests assessing data reliability from three perspectives: data quality, data reconciliation, and data drift. Data quality involves ensuring that data meets certain standards, while data reconciliation checks if data can be traced back to its sources. Data drift refers to changes in the structure or distribution of data over time. The text proposes a framework for data reliability that automates these processes and integrates them into an overall data observability system.
Mar 03, 2022 1,483 words in the original blog post.