Home / Companies / dbt / Blog / August 2026

August 2026 Summaries

9 posts from dbt

Filter
Month: Year:
Post Summaries Back to Blog
Many agentic AI initiatives remain stuck in pilot stages because agents often lack reliable access to governed data, with surveys indicating limited enterprise-scale deployment, widespread data-access constraints, and concerns about trust and governance. Without sufficient context, agents may produce valid but incorrect SQL, misinterpret or invent metric definitions, operate without guardrails or audit trails, and generate excessive compute costs through inefficient processing. The proposed solution is machine-readable data governance rather than manual documentation, using data contracts to define and validate datasets, automated tests to maintain data quality, and a semantic layer to centralize metric definitions and lineage. As AI agents increasingly become major consumers of organizational data, dbt positions its platform as infrastructure for producing trusted, governed data that can support more accurate and scalable AI-driven decisions.
Aug 27, 2026 856 words in the original blog post.
As organizations embed AI and autonomous agents into daily operations, scaling deployment is proving easier than ensuring reliable, explainable, and trustworthy results. The article argues that AI maturity depends on trusted data infrastructure encompassing data quality, governance, clear ownership, business context, interoperability, and cost-efficient operations, rather than on model capability alone. Poor data quality, ambiguous responsibility, unmanaged costs, and difficulty explaining AI outputs become more severe as systems influence more decisions and users; accordingly, 53% of surveyed organizations identify data quality as a major challenge, 41% cite unclear ownership, and 71% worry about incorrect or hallucinated data reaching stakeholders. dbt Labs positions its forthcoming Enterprise AI Data Maturity Model as a five-stage framework for evaluating organizational capabilities, identifying gaps, and prioritizing investments needed to progress from trusted data toward dependable AI and agentic workflows.
Aug 25, 2026 755 words in the original blog post.
A dbt blog post argues that enterprises should avoid using Gong’s transactional API as a high-volume AI context source and instead ingest call data into a warehouse, model and compress it with dbt and warehouse-native AI functions, and expose the results through a dbt MCP server. Direct retrieval of dozens of raw transcripts can consume roughly 240,000 tokens per account query and rapidly exhaust Gong API limits, whereas structured summaries can reduce transcript volume by 20 times or more, with a 10,000-token call compressed to roughly 500–1,000 tokens. The post estimates that this approach can lower AI context costs substantially, citing annual savings of hundreds of thousands of dollars for large sales teams or always-on agent workflows, while also enabling joins with CRM, billing, product, and other warehouse data. It distinguishes among Gong’s official MCP for quick account briefs, an API wrapper for one-off raw transcripts, and dbt’s warehouse-based MCP for filtered, aggregated, joined, trend-based, and governed analytical questions. Although the approach requires upfront modeling work, the author presents it as a reusable pattern for other durable, token-heavy sources such as Slack logs, emails, support tickets, contracts, and marketing materials.
Aug 17, 2026 1,696 words in the original blog post.
dbt Core v1.12 is generally available, introducing framework, Semantic Layer, adapter, and usability improvements while providing an opt-in transition path toward the Rust-based dbt Core v2.0. Its new `--use-v2-parser` flag lets teams test the forthcoming Rust parser, which can be 5–10 times faster on large projects, without replacing the existing Python parser by default. New capabilities include configurable downstream behavior after upstream failures through `on_error`, a separate `vars.yml` file, ad hoc SQL execution via `dbt run-operation --sql`, and expanded UDF support including JavaScript functions on Snowflake and BigQuery, Python functions on Databricks, overloads, and Python package dependencies. The release also updates Semantic Layer definitions by nesting semantic metadata within models and columns, adds Apache Ossie document support, and delivers platform-specific enhancements for Snowflake, BigQuery, Redshift, and Databricks. Other additions include native private-package installation, simplified Iceberg configuration, composable selectors, clearer errors, and automatic pointer views for versioned models, with the release positioned as a gradual preparation step for dbt Core v2.0.
Aug 17, 2026 1,828 words in the original blog post.
Databricks argues that choosing its platform for compute, storage, machine learning, and AI workloads should be considered separately from deciding where transformation logic and data definitions reside. The post contends that placing transformations in Databricks-native tools such as Lakeflow pipelines and notebooks can increase migration costs and make auditing, version control, and metric governance more dependent on platform access, although it acknowledges Unity Catalog’s lineage and governance capabilities. It presents dbt as a complementary transformation layer that stores SQL models, tests, contracts, semantic definitions, and lineage in version-controlled, portable code capable of running across Databricks and other data platforms. The author highlights dbt’s open-source Fusion runtime, cross-platform support, Semantic Layer, community adoption, and state-based incremental builds as reasons organizations may retain independence while using Databricks infrastructure. It recommends that executives assess workload portability, audit access, ownership of metric definitions, and contingency plans before consolidating transformation tooling within a single vendor.
Aug 17, 2026 1,251 words in the original blog post.
Agentic AI projects can produce substantial operational gains, such as automating support and order-processing tasks, but many fail because organizations lack accurate, governed, and accessible data rather than because their underlying models are inadequate. The text identifies operational correctness, control-plane, and human-supervision risks, arguing that agents can amplify errors when they act automatically on stale data, insufficient context, weak permissions, or poorly auditable systems. Successful deployments require centralized and modeled data, lineage, semantic definitions, governance, and interoperability so agents can use authoritative information consistently across systems. Appropriate use cases are high-volume, structured, text- or code-heavy workflows with clear outcomes, low-cost review, and reversible actions, while ambiguous, high-stakes, adversarial, or irreversible work is less suitable. Rather than building models from scratch, organizations can augment foundation models with proprietary data through retrieval-augmented generation and introduce autonomy gradually, beginning with read-only retrieval, then human-reviewed drafting, and finally tightly bounded write actions with approvals for higher-risk decisions.
Aug 14, 2026 1,648 words in the original blog post.
dbt State is a state-aware capability designed to reduce warehouse compute costs and accelerate dbt runs by rebuilding only models whose code or upstream data has changed, while skipping or cloning unchanged models from other environments. Evolving from state-aware orchestration, it is available through the dbt platform, dbt Core v2.0, and a plugin for supported v1 releases, allowing use across production, development, CI, and external orchestrators such as Airflow and Dagster. It tracks hashes of model code and data states in a control plane, then determines whether each model must be rebuilt or can be reused; dbt reports average warehouse-compute reductions of about 30%. Pricing is based on unique daily reused models or tests, termed daily active target tables. Fanatics Betting and Gaming used dbt State with Snowflake, dbt, and Airflow to address overscheduling across thousands of models, increasing reuse from roughly 0.2% to 15% after defining source freshness and model-level lag tolerances, with some projects reaching 25% reuse and an overall average near 8%. Beyond savings, the company reported that the approach simplifies operations by shifting scheduling decisions from job cadence and selectors toward explicit model freshness requirements.
Aug 14, 2026 1,797 words in the original blog post.
dbt Summit 2026, scheduled for September 15–18 at The Cosmopolitan in Las Vegas, will focus on helping data teams prepare governed, scalable data foundations for AI agents and increasingly autonomous analytics workflows. Its keynotes will address improvements to dbt’s engine, AI-ready structured context, AI-assisted development through dbt Wizard, and open infrastructure, while a community keynote will cover the continuing development of dbt Core and its open ecosystem. Product sessions will demonstrate dbt State’s approach to reducing unnecessary compute, dbt Wizard’s development, analysis, and discovery capabilities, and the planned rebuild of dbt Core around a Rust-based Apache 2.0 engine, ADBC adapters, and improved documentation. Other sessions will examine the evolving role of semantic layers and MetricFlow in providing reliable agent context, Apache Ossie’s portable metric specification, open lakehouse infrastructure based on Iceberg, lineage integration between Fivetran and dbt, and methods for supplying agents with governed metadata, semantic models, and usage context.
Aug 10, 2026 1,301 words in the original blog post.
dbt Labs argues that enterprises can reduce the cost and improve the reliability of AI agents by moving raw structured, semi-structured, and unstructured data from vendor MCP connections into a governed context layer in their data warehouse. Drawing on its experience with Gong call transcripts, the company says that directly querying raw data through AI tools created high token and API costs, while ingesting, summarizing, and modeling the same data in the warehouse reduced transcript volume by 20 times and token costs by roughly 98%. The post describes “context engineering” as an extension of analytics engineering in which data teams prepare business context for AI agents rather than only metrics for dashboards, incorporating sources such as emails, support tickets, PDFs, chat logs, and recordings. It recommends compressing, enriching, describing, and governing data; processing expensive LLM reads once in batch jobs; maintaining context incrementally; and providing a shared, reusable context layer accessible to multiple agents. The proposed approach uses Fivetran to move and index data and dbt to model, orchestrate, and apply warehouse-native AI functions, positioning data teams as central owners of trusted AI context without requiring a major change to their existing data stack.
Aug 06, 2026 1,902 words in the original blog post.