Home / Companies / Fivetran / Blog / September 2020

September 2020 Summaries

22 posts from Fivetran

Filter
Month: Year:
Post Summaries Back to Blog
Consistent and accurate data is essential for business operations, marketing, and sales success. The hub-and-spoke data synchronization model is presented as superior to the point-to-point method, particularly in terms of scalability, accuracy, maintenance, security, and overall cost. The hub-and-spoke system connects each app to a central hub, reducing the complexity and number of integrations needed as businesses grow, in contrast to the exponential connections required by point-to-point architectures. This centralization ensures data consistency across the organization, mitigates the risk of data conflicts and overwriting, and simplifies updates and maintenance. Additionally, it enhances security by limiting exposure to data breaches and providing clearer access control. While point-to-point may appear cheaper for individual projects, its long-term operational costs can become prohibitive. Modern data warehouses like Snowflake and BigQuery have made hub-and-spoke systems more affordable, further solidifying this model as the preferred choice for efficient and secure data management.
Sep 29, 2020 1,676 words in the original blog post.
The latest dbt (data build tool) package for Shopify helps businesses understand their customers' purchasing habits better. It generates tables that group customers into cohorts, enriches customer data with additional fields, and adds dimensions to orders, order lines, and product tables. The Shopify API presents challenges in fetching and analyzing data, but Fivetran's Shopify connector simplifies the process by scheduling and batching API calls, unnesting fields, and providing insights into customer and store behavior through dbt packages.
Sep 28, 2020 287 words in the original blog post.
The Modern Data Stack Boot Camp is an educational series that teaches organizations how to implement a modern, cloud-based data architecture in three steps. It covers the components of a modern data stack, the pros and cons of building in-house vs. buying a solution, calculating total cost of ownership for a modern cloud architecture, and choosing the right data stack for an organization. The course is designed to be approachable for beginners while still providing value for those with expertise in data and analytics. Enrollment includes access to three lessons delivered over three days, each focusing on different aspects of establishing a successful modern data stack.
Sep 28, 2020 386 words in the original blog post.
Query planning and optimization is crucial for efficient database performance. SQL, a declarative language, allows users to describe the desired output rather than providing direct execution instructions. When executing a query, the database generates a plan that outlines the steps it will follow to return results. The query planner, a complex component of modern databases, creates this plan. Reviewing a query plan can help identify areas for optimization, such as predicate pushdown and adding indexes. Additionally, data distribution in distributed databases impacts query efficiency, with more efficient joins possible when data is distributed based on relevant join conditions.
Sep 24, 2020 1,804 words in the original blog post.
In 2020, media buying agencies face challenges in integrating data from multiple platforms and differentiating themselves under margin pressure. Modern digital agencies have their mission-critical data stored across various systems, making it difficult to get a comprehensive view of the brand's performance. Additionally, these agencies are under pressure to provide custom solutions that address clients' specific needs. To tackle this issue, Latticework Insights partners with best-in-class technology providers like Snowflake, Tableau, and Fivetran to help agencies build their tech stacks and make data-driven decisions. The partnership between Fivetran and Latticework Insights helps clients centralize multiple data sources into a data warehouse, ready for analysis and reporting.
Sep 23, 2020 570 words in the original blog post.
Fivetran has introduced Fivetran Transformations for dbt Core, allowing users to run their transformation models directly from the application using its native integration. The new feature aims to make data as easily accessible and reliable as electricity. Fivetran's customers needed more advanced features such as SQL transformations, version controlling, peer reviews, testing capabilities, and documentation generation. To address these requirements, Fivetran decided to integrate with the popular open-source software dbt Core by dbt Labs. The integration enables customers to take advantage of a best-in-class automated cloud data integration experience in a single environment. With Fivetran Transformations for dbt Core, users can orchestrate cleaning, testing, transformation, modeling, and documentation of their datasets.
Sep 21, 2020 1,153 words in the original blog post.
Enterprises are facing challenges in unified data analysis despite advancements in marketing. With multiple channels and tools, marketers need to analyze data across platforms to understand which campaigns are successful. Automated data integration can help by routing all marketing data into a centralized place for easy visualization and analysis. Fivetran is a leader in automated data integration with over 150 connectors, providing always-on access to data.
Sep 17, 2020 531 words in the original blog post.
Fivetran, a tool that automates the process of building and maintaining data pipelines, is argued not to eliminate the need for data engineers but rather free them up for more interesting, mission-critical projects. Despite its labor-saving nature, there remains a shortage of data engineering talent and demand for the role continues to grow. Automating ETL allows data engineers to focus on developing proprietary products and niche solutions to help businesses excel. Companies such as Sendbird, The Ignition Group, Square, Fountain, and Billie have benefited from this shift in focus, enabling them to pursue machine learning, improve documentation, and optimize costs.
Sep 16, 2020 852 words in the original blog post.
Fivetran has introduced a new dbt (data build tool) package for Zendesk Support, designed to refine customer success data modeling and enhance detailed ticket tracking. The package provides models that attach response and resolution times, create history tracking for Zendesk ticket fields, and track SLA breaches as per the configured policy in Zendesk. It also supports both business and calendar hour reporting. Fivetran's native Zendesk Support connector helps address challenges associated with the Zendesk API by normalizing data, preserving relationships between tickets and their associated fields, capturing deleted objects, and applying SLA policies to analytics dashboards. The dbt package further enhances this integrated data by offering additional insights into ticket field history tracking and alerting when SLA policies have been breached.
Sep 15, 2020 263 words in the original blog post.
Fivetran has conducted a benchmark comparing the performance and pricing of four popular cloud data warehouses: Amazon Redshift, Snowflake, Presto, and Google BigQuery. The study found that all warehouses had excellent execution speeds suitable for ad hoc, interactive querying. However, each warehouse has unique user experiences and pricing models, with some being more fully managed than others. Fivetran's benchmark results differ from previous studies due to various factors such as data scale, configuration settings, and the complexity of queries used. The most important differences between warehouses are their design choices, emphasizing tunability or ease of use.
Sep 12, 2020 1,891 words in the original blog post.
Fivetran Inc. has released a benchmark comparing the performance and pricing of four popular cloud data warehouses: Amazon Redshift, Snowflake, Presto, and Google BigQuery. The study found that all warehouses had excellent execution speed suitable for ad hoc, interactive querying. However, significant differences exist in user experience and pricing models. While Redshift and BigQuery have both evolved their user experience to be more similar to Snowflake, the market is converging around two key principles: separation of compute and storage, and flat-rate pricing that can "spike" to handle intermittent workloads. The benchmark results indicate that these warehouses all offer excellent price and performance, with the most important differences being qualitative ones caused by their design choices.
Sep 12, 2020 1,891 words in the original blog post.
Fivetran, a fast-growing startup, is developing processes and frameworks to support its expansion while maintaining focus on the company's vision, principles, values, and objectives. The product team aims to create the most reliable, easy-to-use, and connected data integration product for every organization striving to be data-driven. Key principles include connectors as the core of Fivetran, offering simple, predictable, default choices, ensuring data security, and providing quick access to data. The company values are curiosity, ownership, collaboration, kindness, and integrity. Product managers have clear KPIs to measure their success and drive value for users.
Sep 11, 2020 492 words in the original blog post.
Fivetran now offers an online purchasing option where users can "pay as you go" with a credit card. This new feature allows for quick, easy and flexible trial of the service without the need for a full-year commitment or sales-led contracting process. Users can enjoy the full power of data replication during their free 14-day trial period and start consuming Monthly Active Records (MAR) at the end of the trial with monthly billing in arrears. The online purchasing option provides financial flexibility, enabling users to manage cash flow as needed, upgrade or downgrade plans, cancel self-service plan anytime, and accommodate changes in analytical needs. Additionally, all plans purchased through the app qualify for 24/7 customer support.
Sep 08, 2020 582 words in the original blog post.
Indexes are used in databases to speed up queries, similar to how they function in textbooks and libraries. They work by organizing data in a specific order, allowing for efficient searching using algorithms like binary search. While indexes can significantly improve query times, they also have drawbacks such as increased storage space usage and slower write/update operations. Therefore, it's crucial to carefully consider the use cases before deciding to implement an index or not.
Sep 03, 2020 1,463 words in the original blog post.
Consensus refers to broad agreement on the truth in the context of distributed databases, ensuring that all nodes agree on the same values and provide consistent results to queries. The Two Generals Problem illustrates the difficulty of coordinating an agreement between two parties when there's a faulty communication channel. In distributed databases, consensus is generally a problem when multiple nodes operate simultaneously with unreliable network connections between them. Key strategies for making consensus happen include having individual nodes "elect" a leader node in charge of coordinating the log of operations and communicating "truth" to all other nodes. Two widely used consensus algorithms are Raft and Paxos, which involve multiple rounds of votes to reach agreement on an agreed-upon truth.
Sep 03, 2020 1,227 words in the original blog post.
Distributed databases offer powerful solutions to complex problems but also introduce new challenges such as network outages and data distribution issues. A leader node is responsible for coordinating the work of follower nodes, distributing queries among them, and compiling results. When a leader node loses contact with a follower node, it must address questions about backup data and data loss. Updating data in distributed databases requires managing transactions and locking across multiple nodes while considering network outages. Hot segment issues arise when there is an imbalanced distribution of data access across nodes, which can slow down the system. Data shuffling between nodes can also lead to slower response times. The CAP theorem highlights trade-offs in distributed databases: consistency vs. availability and partition tolerance. In the case of distributed databases, network partitions are a fact of life, leading to latency and consistency trade-offs.
Sep 03, 2020 2,098 words in the original blog post.
Distributed and single-node databases differ in their architecture, functionality, and use cases. Distributed databases consist of multiple computers storing data, while single-node databases run on a single computer. Examples of distributed databases include Google Spanner, Azure Cosmos, Redshift, Snowflake, and BigQuery. Single-node databases include PostgreSQL, MySQL, and SQLite. Distributed databases were developed to address the need for storing large volumes of data, speeding up queries by utilizing multiple computers' computational power simultaneously, and ensuring resiliency in case of hardware or network failures. While bigger and better single-node computers can work up to a certain point, they have limitations in terms of cost, size, and fault tolerance. Distributed databases are made up of clusters consisting of nodes (individual computers). There are two main paradigms for distributed databases: big compute and high availability. Big compute involves splitting or sharding data across different nodes to process queries faster, while high-availability databases duplicate data on each node to ensure fault tolerance. In summary, distributed databases allow for more efficient storage and processing of large amounts of data, as well as increased resilience in the face of hardware or network failures.
Sep 03, 2020 1,385 words in the original blog post.
Michael Kaminsky's blog post explores the intricate relationship between isolation and concurrency in databases, emphasizing the importance of understanding these concepts for database performance. As modern databases need to handle multiple transactions simultaneously, managing race conditions—where simultaneous processes affect the same data—is critical to maintain data integrity. The post explains various race conditions like dirty reads, non-repeatable reads, and phantom reads, which can lead to bugs, and introduces isolation levels as a mechanism to prevent such issues. These levels, ranging from Serializable to Read Uncommitted, balance between data consistency and performance, with higher isolation levels providing more consistency at the cost of slower transactions due to increased locking. Locking is a key strategy to manage concurrency, where databases control access to data to ensure transactions are completed without interference. However, excessive locking can lead to deadlocks and performance issues, especially under high contention when multiple users access the same data. The post highlights the trade-offs between isolation and speed, advising careful consideration of isolation levels and locking strategies to achieve optimal database performance.
Sep 03, 2020 2,026 words in the original blog post.
This blog post delves into transactional databases, focusing on the concept of database transactions and their importance in maintaining data integrity. A transaction is a collection of commands that must all be executed together to ensure successful completion or permanent rollback. The ACID properties (Atomicity, Consistency, Isolation, Durability) are crucial for understanding how transactions work. Atomicity ensures that transactions happen completely or not at all, while consistency maintains valid states and prevents invalid transitions. Isolation prevents transactions from interfering with each other, and durability ensures that committed changes are saved to disk and considered permanent until subsequent modifications occur. The series continues with a closer look at isolation levels and locks in the next chapter.
Sep 03, 2020 1,239 words in the original blog post.
In this blog post by Michael Kaminsky, he discusses the differences in architecture between row- and column-based databases. The way data is stored on a hard drive determines whether it's optimized for transactional or analytical workloads. Row stores are efficient for CRUD operations while column stores are better suited for aggregate functions. Understanding disk storage, where data is organized into blocks, helps in comprehending the differences between row and column stores. In row stores, data is written one row at a time, making it suitable for transactional queries that read and manipulate individual objects. On the other hand, column stores organize data by columns, which makes analytical queries faster as they perform aggregate functions on entire columns.
Sep 03, 2020 783 words in the original blog post.
This blog post by Michael Kaminsky discusses the differences between analytical and transactional databases, which are two major paradigms for working with databases. Analytical workloads involve calculating complex aggregate functions, read-only queries, batch-write loads, and ad-hoc non-routine analyses. Transactional workloads operate on one "object" at a time, perform CRUD operations, manage the state of a database, and support many operations per second with high throughput. Common transactional databases include PostgreSQL, MySQL, Microsoft SQL Server, and Oracle Database. The post emphasizes understanding these different paradigms to make informed choices when selecting database technologies.
Sep 03, 2020 622 words in the original blog post.
Databases are integral to modern computing, providing a means to store and organize data for easy access and analysis. The history of databases began with early computing, evolving significantly in the 1960s and 1970s with the development of relational database management systems and SQL, which enabled data to be stored in tables. As the internet expanded in the 2000s, the demand for handling vast amounts of data led to the development of NoSQL databases and big-data processing technologies. More recent advancements include streaming databases optimized for real-time analysis and application-specific databases for niche use cases. Databases can be categorized into analytical versus transactional, relational versus non-relational, distributed versus single-node, and in-memory versus on-disk, each serving distinct functions and optimized for different use cases. SQL, a declarative programming language, is commonly used to interact with databases, though its implementation varies among systems, with some databases like Redis not using SQL at all. The series promises to explore these topics further, delving into the complexities and functionalities of different database systems.
Sep 03, 2020 1,487 words in the original blog post.