December 2024 Summaries
10 posts from CData
Filter
Month:
Year:
Post Summaries
Back to Blog
CData Sync’s Q4 release adds support for replicating data into and out of Microsoft Fabric’s OneLake, an Azure Data Lake Storage-based data lake intended to provide a shared, continuously updated copy of organizational data. The integration supports Parquet, Avro, and CSV files and enables data ingestion from operational sources such as SAP ERP, NetSuite, Workday, and on-premises systems for use in Power BI, while OneLake can also virtualize Delta Lake-based data from platforms including Snowflake, Databricks, Amazon S3, and Google Cloud Storage. The release also introduces an enhanced continuously running change data capture engine for Oracle and PostgreSQL that captures updates every second while aiming to reduce latency and memory use. Additional updates include performance improvements for Snowflake processing and custom roles that provide more granular controls over user permissions and access to customer-created objects.
Dec 19, 2024
451 words in the original blog post.
Data access management (DAM) is the process of defining, enforcing, and monitoring who can access organizational data, under what conditions, and for what purposes, helping protect sensitive information, support compliance, preserve data integrity, and improve operational efficiency. As data volumes and sources expand across on-premises systems, cloud platforms, applications, and third parties, organizations face challenges including fragmented datasets, unclear ownership, manual permissions, slow approval workflows, and limited visibility into access activity. Effective DAM typically involves inventorying and classifying data, assigning roles and permissions through models such as role-based or attribute-based access control, applying least-privilege principles, using strong authentication and encryption, and maintaining approval workflows, logging, continuous monitoring, and periodic audits. Clear policies, employee training, and automated identity and access management tools can reduce unauthorized access, simplify regulatory reporting, strengthen accountability, and ensure users receive appropriate access without unnecessary delays.
Dec 17, 2024
1,790 words in the original blog post.
Data consolidation combines information from disparate sources such as websites, applications, databases, and devices into a unified repository, often a data warehouse or data lake, to create a consistent source of truth. It can improve data quality, accessibility, consistency, governance, reporting, decision-making, operational efficiency, and cost control by standardizing formats, removing duplicates, and reducing data silos. Common approaches include ETL, ELT, data warehouses, data lakes, data marts, manual custom coding, and data virtualization, with the appropriate choice depending on data volume, structure, use cases, and technical requirements. An effective process generally involves identifying sources, mapping schemas, extracting and transforming data, loading it into a target system, merging records, validating quality, and selecting centralized storage based on performance, retention, user, and budget needs. The text also identifies tools including CData Sync, Airbyte, Fivetran, Rivery, and Stitch, highlighting their use of connectors and automated pipelines to support consolidation across cloud and on-premises systems.
Dec 13, 2024
1,495 words in the original blog post.
Manufacturing data and analytics leaders are prioritizing infrastructure modernization, real-time advanced analytics, and automated data engineering to improve supply chains, predictive maintenance, production optimization, customer reporting, and marketing effectiveness. These efforts are complicated by vendor lock-in, fragmented data across legacy on-premises systems, cloud platforms, enterprise applications, and manufacturing systems, as well as costly custom integration architectures. The piece advocates a decoupled, standardized, and automated data connectivity foundation that separates integrations from operational systems, provides governed self-service access, and uses capabilities such as change data capture and schema change detection to replace bespoke ETL pipelines. It cites Recordati, Repligen, and Manhattan Associates as examples of organizations that reportedly reduced vendor dependency, accelerated cloud data initiatives, automated pipelines, lowered costs, and improved decision-making through CData’s connectivity tools.
Dec 11, 2024
1,257 words in the original blog post.
Apache Arrow is an open-source, language-independent framework for high-performance in-memory analytics and data exchange, built around a standardized columnar memory format that avoids costly serialization and deserialization. Its zero-copy sharing, compact memory layout, SIMD-enabled processing, and compatibility with languages including Python, Java, C++, Rust, and Go can improve query, transformation, streaming, and machine-learning workflow performance while reducing interoperability challenges between systems such as Pandas, Spark, Hadoop, and SQL engines. Arrow is particularly suited to large-scale distributed data processing, real-time analytics, and ML pipelines because it can efficiently process relevant columns, manage datasets in record batches, and support modern CPU and GPU architectures. The text also presents CData Connect AI as a complementary connectivity service that provides standardized SQL and API access across Arrow-based, cloud, and enterprise environments.
Dec 11, 2024
1,147 words in the original blog post.
Enterprise data integration (EDI) unifies information from databases, applications, cloud services, and on-premises systems to eliminate data silos, improve access to consistent information, and support faster, better-informed business decisions. Its core components include APIs for system communication, ETL processes for extracting and standardizing data, data governance for quality, security, and compliance, data warehouses for centralized analytics, middleware for connecting incompatible systems, and master data management for maintaining authoritative records. EDI can improve operational efficiency, reduce costs and errors, enhance customer personalization and service, and strengthen collaboration across departments, with applications in finance, healthcare, manufacturing, marketing, and retail. Organizations may adopt on-premises, cloud-based integration platform as a service, or hybrid approaches depending on their infrastructure, scalability, customization, and data-security needs. The text identifies CData, Airbyte, MuleSoft, and Qlik as notable tools, describing CData as a broad real-time integration platform, Airbyte as an open-source connector-based option, MuleSoft as an API-focused enterprise platform, and Qlik as an analytics and reporting solution.
Dec 11, 2024
2,320 words in the original blog post.
CData Sync expanded its reverse ETL capabilities in 2024 from replicating warehouse data into Salesforce to supporting Microsoft Dynamics 365, the second-largest CRM by market share. The feature enables organizations to transfer data from platforms including Snowflake, SQL Server, Google BigQuery, Amazon Redshift, and PostgreSQL into Dynamics, helping businesses that store data outside Microsoft’s ecosystem avoid vendor lock-in while enriching CRM records. Marketing, sales, and support teams can use consolidated information such as purchase histories, website activity, product usage, support tickets, and service agreements to improve targeting, forecasting, customer assistance, and SLA compliance. By making warehouse data available directly within Dynamics, Sync aims to support self-service reporting for business users while reducing routine data-access requests for data engineering teams, allowing them to focus on more complex work.
Dec 10, 2024
449 words in the original blog post.
CData Foundations 2024’s BI and Reporting sessions emphasized that data professionals often spend far more time on behind-the-scenes “data plumbing,” such as connecting, cleaning, and integrating data, than on creating dashboards and deriving insights. Speakers described reliable, real-time access to diverse data sources as essential for effective business intelligence and for preparing organizations to use AI-driven analytics. CData presented its Connect AI platform as a way to automate data connectivity across more than 250 sources, with Kiwi Partners reporting that connection setup time fell from four to five hours per project to less than one hour. The event also highlighted the continuing role of spreadsheets in data exploration and introduced the free CData Connect Spreadsheets product, which enables live data access and source updates within Excel and Google Sheets.
Dec 10, 2024
858 words in the original blog post.
CData’s latest driver update expands connectivity with new integrations for Jira Assets, Okta, and Greenhouse, enabling SQL-based access to asset management, identity, and recruiting data through BI and analytics tools. General improvements add direct SaaS interface links to query results for platforms such as HubSpot, DocuSign, Monday, and Asana, while streaming support for large file uploads improves performance across connectors including Confluence, BigQuery, Asana, and Xero. Connector-specific additions include Sage Intacct write operations using embedded Web Services credentials and Workday WQL pass-through for faster report querying. The release also enhances functionality for HubSpot, QuickBooks Online, Jira, Slack, Snowflake, Google Analytics, SAP, and Google services, alongside major API updates for platforms including Salesforce, Tableau CRM, Workday, Google Ads, Facebook, Kafka, Instagram, and Yahoo Ads.
Dec 03, 2024
678 words in the original blog post.
CData Foundations 2024 brought together hundreds of data professionals to examine how reliable data connectivity supports analytics, AI, operations, and enterprise decision-making. Speakers emphasized that effective data-driven initiatives depend on clean, accessible, governed, and timely data, with security, flexibility, and real-time access forming the basis of a trustworthy source of truth. Discussions addressed preserving source-level permissions through data virtualization, strengthening accountability through centralized auditing, and protecting sensitive information with masking, anonymization, encryption, key management, and regional data-residency controls. The event also highlighted the need for flexible integration across SaaS, hybrid, and multi-cloud environments, enabling organizations to support multiple use cases without dependence on a single vendor. Presenters stressed that combining historical context with live data can improve the relevance of AI, predictive analytics, and operational decisions, particularly in time-sensitive areas such as fraud detection and customer retention, while recordings of the sessions are available on demand.
Dec 02, 2024
1,380 words in the original blog post.