Home / Companies / Fivetran / Blog / November 2021

November 2021 Summaries

19 posts from Fivetran

Filter
Month: Year:
Post Summaries Back to Blog
Connecting Python with Snowflake enables users to optimize development time, enhance machine learning and linear regression capabilities, and accelerate operational analytics by allowing seamless data retrieval and processing in a Jupyter Notebook environment. The tutorial outlines the necessary software requirements, including having a Snowflake database, user credentials, and familiarity with Python. Users are guided through installing the Snowflake Connector for Python, creating a configuration file with authentication credentials, and establishing a connection to Snowflake through Jupyter Notebook. Once connected, the Pandas library can be used to query the Snowflake database, retrieving results into a Pandas data frame for further analysis. This integration facilitates operational analytics by allowing real-time data movement from the warehouse to various SaaS tools, enhancing data accessibility and usability for customer-facing teams. Operational analytics use cases include sending data to tools like Zendesk, Facebook, and Salesforce, with reverse ETL tooling suggested as a more efficient way to handle data transfers.
Nov 30, 2021 1,027 words in the original blog post.
Throughout his career as a solution architect, Dom Orsini has witnessed significant transformations in the data integration industry, evolving from labor-intensive on-premise systems to the agile, fully-managed SaaS solutions of today. Initially, data projects were cumbersome, requiring manual coding and constant supervision, with traditional ETL processes relying heavily on engineering and IT. The advent of big data and infrastructure-as-a-service (IaaS) brought more speed but also complexity and challenges in maintaining interdependent technologies. The current era, marked by the rise of SaaS applications and automated data pipelines, has revolutionized data integration, allowing companies to extract and utilize data efficiently with minimal configuration. Tools like Fivetran and dbt have streamlined data processes, enabling analysts to focus on insights rather than technical hurdles, thus making the field more accessible and dynamic for data professionals.
Nov 23, 2021 800 words in the original blog post.
Databricks has introduced Partner Connect, an ecosystem of partners that allows companies to integrate data tools for easier ingestion, processing, and analysis in their lakehouse. Fivetran is among the first partners featured in Partner Connect, enabling users to ingest and transform their data without manual configuration. With Fivetran on Databricks Partner Connect, end-users can quickly connect popular apps like Salesforce to the lakehouse for various use cases such as analytics, BI, data science, or machine learning. This integration simplifies data management for the lakehouse, allowing joint customers to focus on insights rather than ETL processes.
Nov 18, 2021 546 words in the original blog post.
A recent global survey by Wakefield Research reveals that when enterprises build their own data pipelines, decision-making and revenue suffer. The study found that data engineers spend nearly half their time building and maintaining these pipelines, costing companies an average of $520,000 per year. Many data and analytics leaders reported unreliable and error-prone data from DIY pipelines, leading to poor decision-making and lost revenue. Additionally, the high opportunity cost of in-house pipeline management means that data engineers have less time for advanced data modeling or sophisticated analysis, potentially impacting business outcomes negatively. Automated ELT may provide better results than DIY data pipelines.
Nov 17, 2021 510 words in the original blog post.
This tutorial provides a comprehensive guide on how to export data from Google BigQuery using Python, with various methods including downloading a CSV from Google Cloud Storage, without cloud storage, and using a reverse ETL tool. It explains the steps to create a Google Cloud service account, set up a bucket in Google Cloud Storage, select and query a dataset from BigQuery, and export the data to a CSV file. The tutorial also details an alternative method to write data directly to a CSV using Python libraries like pyarrow and pandas, and introduces reverse ETL tools like Fivetran Activations for more efficient data movement. The guide is designed for users familiar with Google Cloud and Python, offering detailed instructions and summaries to facilitate the data export process.
Nov 16, 2021 2,080 words in the original blog post.
To design a database schema that delivers maximum utility for users, follow these tips: 1) Use entity-relationship diagrams to visualize underlying data models; 2) Ensure schema flexibility to accommodate changing needs; 3) Understand the purpose of your data and its potential uses; 4) Plan ahead by mocking up sample reports or dashboards; 5) Involve both engineers and data analysts in schema design; 6) Optimize indexing for various query types; 7) Implement clear, consistent naming conventions; 8) Consider data security from the beginning of schema design; 9) Document your schema thoroughly to ensure its longevity.
Nov 16, 2021 942 words in the original blog post.
A well-designed database schema is crucial for efficient data warehousing and retrieval. It serves as a blueprint for organizing and categorizing data structures, making it easier for users to obtain actionable information. Poorly designed schemas can lead to confusion, difficulty in modification and maintenance, and wasted effort. Some common issues include outdated entity-relationship diagrams, inflexible schema designs, lack of pre-planning, inconsistent formatting, poor naming conventions, incorrect key field linking, overindexing, manual data normalization or cleansing, and inadequate documentation. To avoid these pitfalls, it's essential to follow best practices for database schema design.
Nov 16, 2021 921 words in the original blog post.
Fivetran, a company that helps businesses manage their data, has fostered a culture of giving back through employee-led charity campaigns. During the COVID-19 pandemic, employees raised over $35,000 for UNICEF India Humanitarian Relief and supported organizations like Black Girls Code to promote equity in the tech industry. The company encourages its employees to pursue their passions and supports them with matching funds for charitable donations. Fivetran is currently hiring for various positions across departments and regions, inviting individuals who share their values of community support and philanthropy to join their team.
Nov 15, 2021 586 words in the original blog post.
The text outlines three methods for transferring data from Snowflake to Hubspot: downloading a CSV file, utilizing Reverse ETL with Fivetran Activations, and using the SnowSQL command line. The CSV download method is simple but not scalable, making it suitable for one-time tasks. Reverse ETL, which does not require manual CSV handling, offers automation and scalability, making it ideal for regular data transfers. SnowSQL provides more flexibility in CSV downloads but still involves manual handling, similar to the first method. For ongoing data syncing, Reverse ETL is recommended due to its ability to automate and schedule data transfers efficiently, ensuring Hubspot always has the latest data.
Nov 10, 2021 1,122 words in the original blog post.
At the Modern Data Stack Conference 2021, Thomas Cooley Professor of Ethical Leadership at New York University, Jonathan Haidt, discussed how emotion and storytelling can be more persuasive than data. He explained that humans are the "last mile" for data, as they often ignore contradictory information. Our irrational side is like an elephant, difficult to control, while our rational side is a small rider on that elephant. Reason acts as a press secretary, always trying to justify what we want to believe. Haidt suggested using storytelling and speaking to deeply held values to change minds more effectively. He also emphasized the importance of understanding others and being open to discussion and disagreement.
Nov 10, 2021 1,069 words in the original blog post.
Fivetran Activations is a reverse ETL tool designed to synchronize data from data warehouses to operational tools, thereby bridging the gap between data teams and various departments within an organization. Unlike ETL tools, Fivetran Activations focuses on operationalizing data by enabling seamless data flows from a central source of truth to tools like CRM, MAP, and other SaaS applications, enhancing decision-making and operational efficiency across departments. The tool is not a data storage platform or a customer data platform but facilitates high-quality data syncs that support operational analytics. To effectively implement Fivetran Activations, the guide recommends understanding the tool's purpose and capabilities, engaging stakeholders early in the process, and initially focusing on impactful, achievable use cases. This approach ensures that all team members benefit from consistent, accurate data, ultimately saving time for data teams and allowing them to focus on more strategic tasks.
Nov 04, 2021 2,600 words in the original blog post.
Fivetran and HVR, two data replication and integration companies, have joined forces to improve their services. They are focusing on enhancing the customer experience by introducing new features such as an intuitive web interface, role-based access control, REST APIs for orchestration, and support for SAP ERP runtime license compatible capture. The combined company aims to offer better overall support with a global 24/7/365 expert team. They are also working on integrating HVR's high-volume real-time replication technology into the Fivetran.com user experience, supporting more complex use cases and adding native support for data integration topologies. Additionally, they plan to simplify pricing into a single consumption-based model and improve data governance by increasing data visibility and controls. The two companies will eventually merge their code bases to deliver a single managed service offering that meets the security, governance, and operational requirements of sensitive organizations such as healthcare, finance, and government.
Nov 04, 2021 860 words in the original blog post.
The article discusses the importance of understanding data transfer and egress costs across Microsoft Azure, Amazon Web Services (AWS), and Google Cloud Platform (GCP). It highlights that while moving storage from data centers to cloud-based environments can result in significant cost savings, decision makers are also concerned about total cost of ownership (TCO) including data egress pricing. The article explains the concept of data egress, which refers to fees for moving data out of a cloud provider's storage system, and contrasts it with data ingress where there is typically no charge.
Nov 04, 2021 304 words in the original blog post.
The article provides a detailed guide on automating custom Fantasy Football logic using reverse ETL, focusing on integrating data from ESPN's Fantasy Football API with Airtable and leveraging SQL and Fivetran Activations for seamless data synchronization. It outlines a multi-step process where the first step involves querying ESPN's Fantasy API to retrieve and modify league data, followed by configuring Airtable to display current standings and projections. The guide then explains how to use SQL to pull statistics into Airtable via Fivetran Activations models, culminating in the final step of running a Python script to trigger data syncs. By following this method, Fantasy Football enthusiasts can establish a reliable source of truth for their league data while customizing it to suit their specific league rules.
Nov 03, 2021 1,034 words in the original blog post.
Fivetran, a company that values having parents and guardians on their team, has formed an employee resource group (ERG) called Fivetran Parents. The aim of this ERG is to provide support and knowledge sharing among working parents and guardians, especially during the pandemic. The group works closely with HR to add benefits as needed and addresses specific concerns related to childcare, flexible work arrangements, and workplace issues. Fivetran also offers generous parental leave policies and a hybrid work model that provides flexibility for employees with dependents. ERGs play a critical role at Fivetran in creating community and supporting underrepresented voices through regular events and seminars open to all employees.
Nov 03, 2021 633 words in the original blog post.
The Looker API is a valuable tool for managing your Looker environment, allowing programmatic access to various aspects of the platform. This blog provides an overview of how to get started using Python and covers important endpoints and objects such as queries, looks, dashboards, dashboard elements, and merge queries. It also presents several use cases, including finding fields and filters in looks and dashboards, and replacing them across multiple queries. The Looker API can be particularly useful for managing large-scale changes or updates to your data environment.
Nov 03, 2021 2,501 words in the original blog post.
Managed data integration is a crucial solution for enterprises to maximize the value of their data and data engineering teams. The DIY approach to integrating data can consume valuable engineering time, as building custom connections to various data sources can take weeks or months per connector. Additionally, maintaining these pipelines imposes significant burdens on data engineering teams. Managed data integration services offer simplicity, speed, and security by connecting disparate data sources quickly and efficiently without diverting engineering resources or introducing new data security risks. Most enterprises can rely on a managed data integration service for 80-90% of their needs, with the option to build custom connectors when necessary. This hybrid approach allows companies to take advantage of the efficiency and speed of a managed solution while maintaining flexibility for future data integration options.
Nov 02, 2021 1,222 words in the original blog post.
Marketing analytics is the process of evaluating marketing initiatives' success and value, identifying trends and patterns over time, and making data-driven decisions. It helps measure campaign performance, find opportunities in marketing performance, understand customers, and track competition. The marketing funnel consists of awareness, consideration/conversion, and retention stages. To ensure data analytics success in marketing, marketers should access combined data from multiple sources, build meaningful marketing dashboards for specific use cases, choose appropriate analytics visualizations, develop models that measure and predict, and analyze past performance to improve future results. Automated data pipelines can facilitate precise targeting, accurate lead attribution, and maximal campaign ROI.
Nov 02, 2021 975 words in the original blog post.
An enterprise data warehouse (EDW) is a centralized repository for storing and analyzing business data from various departments and sources. It enables data-driven decisions, quicker time to insight, and consolidated and standardized data across an organization. Cloud-based EDWs offer benefits such as speed, scalability, lower total cost of ownership (TCO), cloud elasticity, integration capabilities, and self-service for business users. Key selection criteria when choosing a cloud data warehouse include compatibility with existing systems, cost comparison, scaling capacity, security features, user access control, fault tolerance, and peer recommendations.
Nov 01, 2021 1,560 words in the original blog post.