December 2018 Summaries
15 posts from DataStax
Filter
Month:
Year:
Post Summaries
Back to Blog
DataStax has released version 6.7 of its DataStax Enterprise (DSE) platform, which includes new enhancements from technology partners. Key improvements include the DSE Metrics Collector for integrating with standard centralized monitoring solutions, production-ready Docker images for DSE, Studio, and OpsCenter, and the DataStax Apache Kafka Connector for improved integration between DSE and Apache Kafka. These enhancements aim to provide developer and operator simplicity while addressing customer needs for an extensible platform and ecosystem of complementary technologies.
Dec 20, 2018
488 words in the original blog post.
The latest release of DataStax Enterprise (DSE) introduces enhancements to DSE Analytics, making its features more secure, faster, and easier to use. Key updates include support for Kerberos authentication in the DSEFS REST API, convenience methods for recursive commands, improved logging in the DataStax Spark Shell, and intelligent throttling mechanisms between Apache Cassandra™ and Spark™ components. These improvements aim to provide a better experience for developers working with DSE Analytics and strengthen its usability for analytics users.
Dec 19, 2018
545 words in the original blog post.
Part 3 of the blog series explores the use of DataStax Enterprise Analytics and various technologies like Apache Cassandra, Apache Spark, PySpark, Python, and Jupyter Notebooks for conducting text analytics, specifically sentiment analysis on movie-related Twitter data. The process involves setting up the necessary environment, including DataStax and the Twitter Developer API, then using Apache Spark’s MLlib functions such as Tokenizer and StopWordsRemover to preprocess the tweets. The analysis is performed using the Pattern library in Python to determine sentiment scores from cleaned tweets, which are stored in Cassandra tables. By comparing positive and negative sentiment scores, the analysis concludes whether a movie is generally liked based on the Twitter data, exemplified by an analysis of "SpiderVerse," which was found to be positively rated. The blog emphasizes the iterative nature of data science and encourages further exploration and feedback.
Dec 19, 2018
2,401 words in the original blog post.
DataStax is inviting the Cassandra community to submit papers for its premier conference, DataStax Accelerate, taking place on May 21-23, 2019 in National Harbor, Maryland. The Call for Papers (CFP) closes on February 15, 2019. Speaking at the event can benefit one's career, contribute to a stronger Apache Cassandra community, and provide opportunities for free or discounted conference passes. Tips for preparing an engaging talk include focusing on specific lessons learned, tailoring content to the audience, providing sufficient details in the submission, avoiding lengthy text slides, and refraining from sales pitches.
Dec 18, 2018
842 words in the original blog post.
DataStax Enterprise (DSE) 6.7 is a new version of Apache Cassandra™ that offers improved performance, stability, and reduced complexity for users. It introduces a new search shard selection algorithm to reduce resource usage and enhance the execution of flexible queries. DSE 6.7 also elevates native geospatial functionality, making it easier to configure, execute, and manage geospatial queries and data types. The update supports automatic creation of search indexes by inferring data types in CQL schema and generating corresponding Solr schemas. Additionally, it handles native spatial data types and adds support for Solr's facet heatmap functionality through the CQL solr_query API.
Dec 18, 2018
411 words in the original blog post.
The Node.js Object Mapper for Apache Cassandra has been released, allowing users to interact with data like a set of documents. The Mapper is provided as part of the driver package and can be used across applications. It contains a single ModelMapper instance per model in an application, which retrieves and saves objects from and to the database. The Mapper also supports mapping a single model to multiple tables or views for more efficient reads. Additionally, it introduces an internal mechanism to cache query functions based on object shapes, optimizing performance. This new Object Mapper is available in DataStax Node.js Driver for Apache Cassandra version 4.0 and DataStax Enterprise Driver 2.0.
Dec 16, 2018
1,013 words in the original blog post.
DataStax has released version 4.0 of its Node.js Driver for Apache Cassandra and version 2.0 of the DataStax Enterprise Node.js Driver, introducing new features such as RequestLogger, Object Mapper, improved load-balancing policy, support for JavaScript BigInt type, exposed internal driver metrics, explicit local data center setting, and control over retrying when a timeout/error is encountered. These updates aim to enhance the performance and functionality of the drivers.
Dec 16, 2018
972 words in the original blog post.
DataStax has introduced the DSE Metrics Collector in its latest version, DSE 6.7. This tool allows users to integrate DSE metrics with any standard centralized monitoring solution of their choice, such as Prometheus, Nagios, Graphite, and more. The DSE Metrics Collector is based on an industry-standard collector and comes with a comprehensive set of knobs and configuration options that can be customized based on the user's environment. It also supports automatic management of the service by the DSE server, eliminating the need for separate agent software. Additionally, DataStax has provided a getting-started Github project to automate the setup process and explore sample dashboards.
Dec 13, 2018
601 words in the original blog post.
The DSE OpsCenter Backup Service in version 6.7 introduces significant enhancements to backup and restore processes for DataStax Enterprise clusters. Restore performance has been optimized, resulting in restores being up to six times faster than previous versions. Additionally, support for S3-compatible object storage destinations and Azure Blob Storage as a backup destination have been added, providing more flexibility in storing backup data. These improvements aim to enhance the overall quality and reliability of OpsCenter while promoting vendor lock-in avoidance and accelerating hybrid cloud deployments.
Dec 12, 2018
437 words in the original blog post.
DataStax has made the Apache Kafka® Connector freely available for Open Source Cassandra users. The DataStax Apache Kafka Connector, built by the team that authors the DataStax Drivers for Apache Cassandra®, seamlessly moves data from Apache Kafka to DataStax Enterprise (DSE) in event-driven architectures. It offers market-leading performance, flexibility, security, and visibility. The connector is fully supported by DataStax and provides expert services. Key features include consuming Kafka primitive, JSON, and Avro data formats; working with various connectors; providing JMX metrics; running within Connect Worker; offering at least once guarantee for records; supporting standalone mode and distributed mode/HA support; enabling flexible Kafka topic to DSE table mapping; allowing a single Kafka topic to be written to multiple DSE tables; having built-in throttling and parallelism; accounting for varying date/time formats; configuring consistency levels, row-level TTLs, deletes, null handling, error handling, offset management, SSL connections, LDAP/Active Directory, Kerberos, write timeouts, and compression strategies.
Dec 11, 2018
1,066 words in the original blog post.
DataStax has joined Microsoft's Private Offers program, allowing them to create exclusive offers for their closest customers on Azure Marketplace. This move is in response to the growing demand for convenient and efficient software distribution channels, similar to those found in consumer app stores like App Store and Google Play. The private offerings will enable users to pay-as-you-consume, click-to-accept terms, and consolidated billing onto their Azure bill.
Dec 10, 2018
413 words in the original blog post.
This tutorial demonstrates how to enable a new DSE Metrics Collector demo using a DataStax Enterprise (DSE) 6.7 Docker container, export metrics to a Prometheus Docker container, and visualize these metrics with a Grafana Docker container. The process involves cloning the DataStax DSE Metrics Reporter GitHub repo, navigating to the demo directory, starting containers using docker-compose, and accessing the metrics through provided Grafana dashboards. To remove the containers, simply run "docker-compose down".
Dec 10, 2018
272 words in the original blog post.
DataStax has released DSE 6.7, along with updates to OpsCenter, DataStax Studio developer tool, and DataStax drivers. The latest version includes enhancements to DSE Analytics, improved geospatial search capabilities in DSE Search, optimized restore engine, support for backups to Microsoft Azure and generic S3 platforms, first phase of the DSE Insights engine, and an Apache Kafka Connector. Additionally, certified Docker images are now available for production use.
Dec 05, 2018
739 words in the original blog post.
As a Solutions Engineer at DataStax, the question of how to confirm correct data loading in a DataStax Enterprise (DSE) cluster is frequently asked. This is particularly crucial when data governance is important or DSE becomes the System of Record. Traditional databases have various methods for reconciling data between environments, but this can be more challenging with Apache Cassandra due to its distributed nature. However, there are options available:
1. AlwaysOn SQL in DSE 6.0 allows for highly-available and secure SQL service execution directly in Studio.
2. SEARCH INDEX creation enables validation of date fields and checking for blank fields.
3. DSE Search CQL Sum and Cassandra Count can be used to return the total number of rows, validate date fields, and check field values.
4. DSE Analytics with Apache Spark integration is recommended for identifying discrepancies between data sources.
5. Monitoring application logs and system.log files on each node, as well as using OpsCenter or nodetool tablestats, can help detect any issues impacting reconciliation queries.
Dec 04, 2018
381 words in the original blog post.
Developer Day in Paris brought together 70 developers to learn about Cassandra features, use cases, and strengthen their knowledge on specific topics. The event aimed at improving participants' understanding of technology for better internal application. Some companies like Capgemini attended the workshops covering all topics. Developers expressed interest in more technical challenges and some regretted having to choose between different workshops. Overall, the event was beneficial for developers looking to improve their knowledge on Cassandra and DataStax solutions.
Dec 03, 2018
599 words in the original blog post.