January 2025 Summaries
10 posts from CData
Filter
Month:
Year:
Post Summaries
Back to Blog
CData Connect AI plans to introduce scheduled queries at the end of January, enabling users to automatically run Query Builder or custom SQL queries from the Data Explorer and write results to selected destinations at defined intervals. The feature is intended to support low-maintenance data workflows without requiring custom scripts or complex pipeline infrastructure, helping users keep reporting, analytics, and project-management data current. Supported destinations include BigQuery, Databricks, PostgreSQL, Redshift, Smartsheet, and Snowflake, while the platform’s broad source connectivity allows data to be consolidated and delivered across systems. Users create or write a query, choose a schedule and destination, and Connect AI executes the data write automatically; the capability will be available to all users, with a free trial offered through the Data Explorer.
Jan 30, 2025
345 words in the original blog post.
Retailers such as Amazon, Walmart, Nordstrom, and Costco have used data infrastructure and predictive analytics to improve demand forecasting, supply chains, marketing decisions, and customer experiences, illustrated by Amazon’s response to the pandemic-era toilet paper shortage. The passage argues that many organizations still struggle to unify data across expanding mixes of legacy, cloud, on-premises, and vendor-specific systems, limiting their ability to support analytics and artificial intelligence. It identifies three core requirements for modern retail data operations: ecosystem-agnostic architecture that connects disparate systems, a central governed data-access layer that makes information available to analytics teams, and automated data engineering that reduces manual work and technical debt. Examples involving Office Depot and BJ’s Wholesale are presented to show how cross-platform integration and automated data preparation can improve performance, business continuity, workforce analysis, and retention.
Jan 27, 2025
1,431 words in the original blog post.
CData’s Foundations conference highlighted how Google BigQuery, UiPath, APOS Systems, Klipfolio, and Stellar One embed CData’s connector library to broaden data access, reduce the burden of building and maintaining integrations, and support new product capabilities. Google described connectors as a way to unlock enterprise and legacy data for transformation, reporting, and generative AI, noting that reliable AI depends on accessible, high-quality data. UiPath uses CData to extract high-volume process data from systems such as SAP, Salesforce, and ServiceNow for its Process Mining product, with roughly 75% of its customers licensing the technology. APOS integrated CData JDBC drivers into its SAP Analytics Cloud gateway to provide live access to non-SAP sources, while Klipfolio replaced a resource-intensive in-house connector effort with a white-labeled, scalable connectivity layer for self-service analytics. Stellar One adopted CData’s hosted Connect AI platform to simplify data migration during Acumatica ERP implementations, illustrating how outsourced connectivity can help smaller providers address complex integration needs.
Jan 21, 2025
1,482 words in the original blog post.
ETL, or extract, transform, load, is a data integration process that gathers information from sources such as databases, APIs, files, and cloud services, cleans and restructures it to meet business and analytical requirements, and loads it into centralized destinations including data warehouses, cloud storage, and analytics platforms. It supports reliable reporting, business intelligence, machine learning, cloud migrations, IoT analytics, database replication, and industry-specific applications by improving data accessibility, quality, security, scalability, and operational efficiency while reducing manual work, errors, and storage costs. Organizations can run ETL in batch, streaming, or incremental modes and select from tools such as CData Sync, Airbyte, Apache Airflow, AWS Glue, Azure Data Factory, Google Cloud Dataflow, Informatica, Matillion, Microsoft SSIS, Talend, and others based on their infrastructure and integration needs. CData Sync is presented as a platform offering prebuilt connectors and real-time synchronization between on-premises and cloud systems to help create analysis-ready data pipelines.
Jan 20, 2025
1,652 words in the original blog post.
Hybrid integration platforms (HIPs) and integration platforms as a service (iPaaS) are two approaches for connecting the growing number of business systems that collect and exchange operational data. HIPs support both on-premises and cloud deployments, making them suitable for organizations that must integrate legacy infrastructure with cloud applications; they offer flexible deployment, centralized architecture, security, and agility, but can require greater technical expertise, more complex tooling, and higher planning costs for maintenance and scaling. iPaaS platforms are cloud-first services designed primarily for cloud-to-cloud integration, providing lower upfront costs, rapid deployment, prebuilt connectors, simplified administration, and easier scalability, although they have limited support for on-premises data transformation and connectivity and may face performance bottlenecks with large or complex workloads. Choosing between them depends on an organization’s existing systems, integration complexity, IT skills, budget, and growth plans: HIP is generally better for hybrid environments with substantial on-premises requirements, while iPaaS is better suited to cloud-focused businesses seeking a scalable, lower-maintenance integration solution.
Jan 16, 2025
1,405 words in the original blog post.
Operational databases and data warehouses serve complementary but distinct roles in data management: operational databases support real-time, day-to-day transactions, while data warehouses organize historical information for analytics and strategic decision-making. Operational systems use OLTP, row-oriented storage, normalized tables, and ACID transaction controls to efficiently process frequent inserts, updates, and simple queries for applications such as e-commerce, CRM, point-of-sale, and logistics. Data warehouses use OLAP, column-oriented storage, denormalized star or snowflake schemas, and scheduled batch updates to support complex queries, reporting, trend analysis, marketing analytics, and financial forecasting. The appropriate choice depends on whether an organization needs live operational data or scalable historical analysis, though many businesses use both systems together and integrate data sources through tools such as CData Sync.
Jan 14, 2025
1,259 words in the original blog post.
Apache Iceberg and Delta Lake are open table formats that add structure, governance, scalability, and reliable data management capabilities to data lakes, supporting analytics, AI/ML, and large-scale processing workloads. Iceberg, created by Netflix and maintained as an Apache project, emphasizes multi-engine compatibility, scalable Parquet-based metadata, hidden partitioning, schema evolution, atomic operations, and efficient handling of large, read-heavy or batch-oriented datasets. Delta Lake, initiated by Databricks, is closely integrated with Apache Spark and provides ACID transactions, transaction-log metadata, file compaction, indexing, and time travel, making it particularly suitable for write-intensive, real-time, streaming, machine learning, and data warehousing applications. Both formats provide versioning, rollback or historical query capabilities, and storage optimization services, but Iceberg is positioned as more flexible across varied engines and cloud infrastructures, while Delta Lake is especially advantageous in Spark- and Databricks-centered environments requiring strong consistency during frequent updates. Selecting between them depends on data consistency needs, existing tools, workload patterns, retention requirements, and desired ecosystem flexibility, while platforms such as CData Connect AI aim to provide standardized access across these and other data lake environments.
Jan 10, 2025
1,112 words in the original blog post.
Structured data is information organized into predefined formats, typically rows and columns in databases or spreadsheets, that makes it easy to store, search, query, analyze, and automate with tools such as SQL and business intelligence platforms. Its consistent schema, relational links, and defined fields support efficient reporting, forecasting, AI and machine learning applications, and operational systems such as customer relationship management, financial transaction processing, reservations, inventory management, marketing analytics, and electronic health records. Common tools include relational databases such as MySQL, PostgreSQL, and SQLite, along with OLAP platforms and spreadsheets. Compared with unstructured data such as videos and social media posts, and semi-structured formats such as JSON and XML, structured data offers greater accessibility, storage efficiency, scalability, and analytical simplicity, but its rigid schemas can limit flexibility and make adaptation to changing requirements or dynamic data types costly.
Jan 08, 2025
1,538 words in the original blog post.
CData Foundations 2024’s Data Architecture Track featured leaders from PartnerRe, WashTec, and NYU discussing how modern data platforms can improve collaboration, agility, governance, and AI readiness. PartnerRe emphasized breaking down business and IT silos through shared rules, democratized data access, and a scalable platform supporting compliance and analytics. WashTec described using a data mesh approach and robust infrastructure to let teams experiment and pursue AI initiatives without managing excessive technical complexity. NYU shared its incremental five-year adoption of data virtualization to unify distributed institutional data, build stakeholder alignment, and improve data literacy. Across the sessions, speakers identified accessible, governed, scalable, and reusable data products as essential foundations for timely insights and meaningful AI-driven innovation.
Jan 07, 2025
621 words in the original blog post.
CData has introduced drivers and connectors for Salesforce Data Cloud that enable live integration with more than 200 BI, analytics, AI, reporting, and data integration tools, reducing the technical complexity of accessing the platform through its API. Salesforce Data Cloud consolidates Salesforce and external SaaS, CRM, marketing, support, and database data to support customer 360 initiatives and deeper analysis. The new offerings include JDBC, ODBC, and ADO.NET connectivity, native connectors for Excel, Tableau, and Power BI, plus Python and PowerShell options, with CData Sync and CData Connect AI planned for future releases. They support SQL-92 queries, several authentication methods including OAuth and JWT, and access to Calculated Insights, Data Graphs, Data Lake Objects, and Data Model Objects. A Power BI example illustrates installation, OAuth-based connection setup, and use of DirectQuery to select and visualize Salesforce Data Cloud data without coding.
Jan 07, 2025
713 words in the original blog post.