Home / Companies / Fivetran / Blog / August 2026

August 2026 Summaries

11 posts from Fivetran

Filter
Month: Year:
Post Summaries Back to Blog
Fivetran’s platform engineering leaders describe how AI changed performance work by sharply reducing the time required to gather context from production logs, code history, telemetry, and complex execution paths, making it economical to evaluate many previously overlooked incremental improvements. The team used AI for codebase archaeology, static analysis of hot loops, fleet-wide log mining, rapid prototypes and microbenchmarks, reusable diagnostic skills, and automated operational tooling, while emphasizing that AI-generated findings require validation through trusted metrics and benchmarks. Reported outcomes included up to 70% benchmark import-throughput gains from asynchronous BATCH_COMPLETE signaling, roughly doubled SAP HANA benchmark throughput, 35–40% SAP HANA production improvements, elimination of over 90% of API calls in some GitHub connector scenarios, and targeted connector improvements such as parallelized Okta queries. The authors argue that AI’s primary benefit is not faster coding but faster identification and sizing of worthwhile work, enabling teams to inspect complete fleets rather than samples, avoid low-value projects through negative findings, and automate useful but historically unprioritized tasks. They conclude that AI is most effective when paired with established measurement infrastructure, performance benchmarks, and human judgment for prioritization.
Aug 28, 2026 3,096 words in the original blog post.
Chief Data Officers face expanding expectations in the AI era, with their role centered on ensuring data access, creating valuable data products, and governing data responsibly, while the average tenure remains about 30 months. The piece argues that unclear mandates, overlapping responsibilities with other technology executives, and poorly defined measures of success often create more difficulty than technical challenges. It recommends explicitly linking the CDO mandate to business-focused OKRs and KPIs covering AI-ready data coverage, delivery speed, product adoption and value, and governance outcomes. Rather than pursuing a rigid or comprehensive infrastructure transformation, organizations should begin with a high-value use case and iteratively develop reusable capabilities in data cataloging, ownership, access control, integration, modeling, observability, semantic context, and open, interoperable data infrastructure. As AI systems require explicit, reliable context that human workers may infer from experience, centralized and well-governed data becomes increasingly important, and proposed AI initiatives should be evaluated for technical feasibility, economic viability, and legal, social, and organizational acceptability.
Aug 26, 2026 1,186 words in the original blog post.
Fivetran describes how it substantially improved data-pipeline performance after its existing incremental optimization efforts failed to meet an enterprise target of moving 1 TB of data in two hours. Starting from benchmarks of 26 MB/s for Oracle HVA-to-Snowflake and 48 MB/s for Postgres-to-BigQuery, the company ultimately reached 139 MB/s and 259 MB/s respectively, while increasing core production throughput from 7 MB/s to 70 MB/s. Its approach centered on establishing serialized extract volume as a consistent throughput metric, instrumenting every pipeline phase, creating repeatable benchmarks with controlled environments, building visualization and profiling tools, and documenting both gains and regressions. A cross-functional temporary “Tiger Team” was given protected focus and authority across connector, core, destination, infrastructure, and quality-engineering boundaries, enabling system-wide changes such as parallel processing, multithreaded imports, serialization improvements, and connector rewrites. After meeting its goals through roughly 50 projects, Fivetran disbanded the temporary team and formed a permanent Platform Engineering performance group to maintain benchmarks, prevent regressions, improve efficiency and scalability, and support customer-specific performance needs.
Aug 24, 2026 3,754 words in the original blog post.
Fivetran’s August 2026 discussion of build-versus-buy decisions for data pipelines argues that the main expense of custom integrations is ongoing maintenance rather than initial development, including responding to API and schema changes, outages, security needs, monitoring, and new data-source requests. It cites benchmarks claiming organizations allocate 7.5% of data budgets and 53% of engineering time to pipeline maintenance, while custom or legacy integrations reportedly fail more often than managed alternatives. The piece contrasts this with managed platforms, which use prebuilt connectors, automated maintenance, and configurable transfers to reduce setup time and allow teams to focus on analytics, governance, product work, and AI initiatives. It also emphasizes the opportunity cost of assigning lean engineering teams to infrastructure work, particularly where reliable, current, and complete data is needed for AI projects. Customer examples are presented as evidence of reduced maintenance and labor costs, and the discussion concludes by promoting Fivetran’s more than 750 connectors and Connector SDK as options for centralized, managed data movement, including support for proprietary systems.
Aug 19, 2026 2,059 words in the original blog post.
Fivetran presents data destinations as a central architectural decision because they determine how data is governed, analyzed, modeled, and returned to business systems. Its platform standardizes schemas and metadata across destinations, aiming to make dbt projects, SQL logic, and prebuilt analytics packages portable among warehouses such as Snowflake, BigQuery, Databricks, and Redshift. Tables can use either soft-delete mode for current-state records or history mode for native SCD Type 2 tracking, with automatic restructuring when modes change. Fivetran also supports managed open data lakes using Iceberg and Delta formats in cloud storage, enabling multiple query engines to access shared data while automatically updating catalogs such as AWS Glue, Unity Catalog, and BigLake. Through its Activations reverse ETL service, governed warehouse data can be synchronized to operational tools including CRMs, marketing platforms, advertising systems, and collaboration applications. The company supports more than 30 destination types, spanning cloud warehouses, data lakes, analytical and relational databases, streaming services, and vector databases, with common connector coverage and optional hybrid deployment for data residency needs.
Aug 18, 2026 1,581 words in the original blog post.
dbt Wizard is an AI agent designed for dbt projects that uses full project metadata, including models, lineage, tests, contracts, semantic definitions, and development configurations, to provide context-aware analytics engineering assistance. Available through the dbt platform and dbt Core command line, it can answer project questions and perform actions such as documenting projects, creating models, refactoring fields across files, migrating legacy logic, assessing downstream impacts, and updating semantic-layer metrics. Unlike generic coding assistants, it searches and interprets the dbt DAG, identifies model grain and dependencies, resolves ambiguous references, follows column lineage, and avoids unsafe changes to widely used models. It also validates work through dbt compilation, builds, tests, warehouse queries, and specialized validation agents, aiming to reduce hallucinations, broken models, metric conflicts, and manual review effort while accelerating data product development.
Aug 18, 2026 1,108 words in the original blog post.
dbt Wizard is a conversational AI tool for dbt projects that uses project-specific lineage, compiled state, tests, and semantic definitions to generate and modify data models with greater context than generic SQL assistants. Available as a free CLI or hosted dbt platform feature, it offers Build, Refactor, Investigate, and Migrate modes to create governed models, manage dependencies, diagnose issues, and move models safely. The post argues that the tool can reduce the translation effort between business requirements and data implementation by producing SQL alongside tests, documentation, and metric definitions while validating proposed changes against existing project conventions and downstream dependencies. It is designed to support both new and mature projects, layered data architectures, industry-standard schemas, and data lake modernization workflows. The author positions Wizard as a way for Fivetran partners to accelerate delivery of reliable, agent-ready analytics models, complementing Fivetran’s data-ingestion automation and potentially shortening projects from months to days.
Aug 13, 2026 1,249 words in the original blog post.
Healthcare organizations are adopting AI for clinical documentation, diagnostics, patient monitoring, engagement, research, and revenue-cycle operations, but many face barriers from legacy ETL systems that are slow to build, limited to predefined use cases, and unable to provide continuously updated data. Fivetran and phData argue that an AI-ready foundation should use automated ELT, broad source replication, change data capture, schema-change management, encryption, and governance to make structured and unstructured healthcare data available in a common data layer. The approach combines Fivetran’s connectors, including for Epic and other healthcare platforms, with dbt for tested, documented data models and faster incremental transformations, while Fivetran Activations can return trusted data to operational systems such as CRMs and EMRs. The companies emphasize that this architecture can reduce custom engineering and support compliance requirements, citing a biopharma deployment and Inova Health’s reported acceleration of a planned four-year transformation into six months, and invite readers to a September 22 webinar on healthcare data platforms.
Aug 10, 2026 1,602 words in the original blog post.
Open Data Infrastructure aims to support diverse data use cases through interoperable systems and a unified source of truth, but the article argues that this also requires context engineering to make data understandable, trustworthy, and retrievable for both people and AI. Fivetran and dbt Labs position the Managed Data Lake Service as an approach for continuously synchronizing source data into open table formats, managing schema changes, and publishing governed metadata for use across compute engines. Context engineering addresses four areas: defining data semantics and provenance, validating quality and operational state, governing access and lifecycle changes, and delivering only the relevant machine-readable context to specific users or AI agents. The article highlights dbt’s Semantic Layer, Apache Ossie, dbt Catalog, Discovery API, dbt Mesh, and MCP server as tools and standards for documenting business meaning, lineage, testing, freshness, governance, discovery, and AI integration. By externalizing institutional knowledge and making it explicit, organizations can improve self-service, collaboration, reproducibility, onboarding, and the reliability of decisions and automated systems.
Aug 07, 2026 1,125 words in the original blog post.
Fivetran replaced its heavily customized Atlassian Statuspage deployment, which cost about $65,000 annually plus maintenance effort and suffered from performance, rate-limit, and update-propagation issues, with an internally built status page created by two engineers using AI coding agents. Over roughly four months, the team established detailed requirements, API contracts, design assets, agent roles, reusable instruction libraries, and a structured development process before producing an MVP, demonstrating that extensive specification and human oversight were essential to effective AI-assisted development. Although the initial MVP was completed quickly, production hardening—including authentication, infrastructure, integrations, audit logs, subscriber migration, load testing, and bug fixes—accounted for 56% of the approximately 1,050 engineering hours and brought the estimated one-time cost to about $109,500 including AI usage. The new system now serves roughly 26,000 subscribers, was validated during a real high-severity incident, operates independently from Fivetran’s core platform, and is projected to cost $12,400 to $25,000 per year, yielding a roughly 2.5-year payback period. Fivetran concludes that AI can improve the economics of replacing well-specified, relatively simple SaaS tools with poorly aligned pricing, but it does not eliminate the substantial work of production integration, reliability engineering, and human judgment.
Aug 07, 2026 3,222 words in the original blog post.
Fivetran's Managed Data Lake Service now includes the ability to execute DDL and DML operations, enhancing its flexibility and appeal over traditional data lakes and cloud data warehouses. By decoupling storage from compute, it offers benefits such as reduced ingestion costs and ACID compliance, while maintaining data in open formats like Apache Iceberg™, allowing for interoperability across multiple compute engines like Snowflake and BigQuery. The introduction of Write Credentials enables users to execute operations like deleting or dropping tables, crucial for scenarios such as GDPR compliance, while maintaining control through a separate set of credentials that limit write access to specific users. This feature not only supports migrations to a more open, cost-effective infrastructure but also ensures that data integrity is preserved by advising users to pause connections during data modifications. This advancement opens up more use cases and promotes the adoption of an Open Data Infrastructure, encouraging consideration of upgrading storage solutions for greater flexibility and cost efficiency.
Aug 03, 2026 1,402 words in the original blog post.