Home / Companies / Elastic / Blog / December 2015

December 2015 Summaries

19 posts from Elastic

Filter
Month: Year:
Post Summaries Back to Blog
Exceptionless, an open-source technology company specializing in real-time event reporting and error logging, transitioned from MongoDB to Elasticsearch to improve scalability, efficiency, and ease of use. Initially using MongoDB for its scalability and cross-platform capabilities, Exceptionless faced challenges with real-time dashboards, backups, and storage as they grew. Elasticsearch resolved these issues by providing faster indexing, live statistics across time zones, simpler backup and restoration processes, and reduced disk space usage. The switch also simplified cluster management and configuration, offering a more user-friendly and open-source-friendly experience. The company continues to benefit from Elasticsearch's ongoing enhancements, appreciating its simplicity and openness.
Dec 22, 2015 803 words in the original blog post.
ITV, the UK's oldest commercial TV network, has integrated the Elastic Stack into its operations to modernize its content distribution infrastructure, particularly focusing on non-aerial platforms like the ITV Hub. The challenge of managing disparate logging systems across multiple legacy and new products led ITV to adopt a unified logging solution using the Elastic Stack, which includes Elasticsearch, Kibana, and Logstash. This integration has streamlined log management, enabling the creation of shared dashboards for both development and operations teams, thus enhancing visibility and collaboration. Initially tested and adopted by ITV's Common Platform team, the Elastic Stack is now central to ITV's operations, handling millions of log entries daily and supporting over 11 million users on the ITV Hub. This strategic move has not only improved operational efficiency but also facilitated real-time data visualization and error reporting, reflecting ITV's commitment to leveraging advanced technology for competitive advantage. Efstathios Xagoraris, a Platform Engineer at ITV, has been pivotal in this transformation, advocating for the Elastic Stack's adoption to ensure a scalable and cohesive logging framework across the company.
Dec 21, 2015 951 words in the original blog post.
Kristina Frost reflects on the holiday season as an opportunity to embrace a culture of giving, inspired by her childhood experiences of family gatherings and gift exchanges. As an adult, she finds joy in giving thoughtful presents and emphasizes the importance of extending generosity beyond one's immediate circle to those less fortunate. At Elastic, she participates in philanthropic initiatives supporting underserved communities, particularly in STEM education and food security. The company's global efforts include raising substantial funds for charities and organizing local events to provide gifts and essentials to children in need. This spirit of giving fosters a sense of community and highlights the impact of collective small acts of kindness, underscoring the belief that while individuals cannot do everything, they can still contribute meaningfully to make a difference.
Dec 21, 2015 1,363 words in the original blog post.
On December 17, 2015, Elasticsearch announced the release of versions 2.1.1, 2.0.2, and 1.7.4, which introduced significant bug fixes and improvements. The 2.1.1 version, based on Lucene 5.3.1, addressed critical issues such as translog corruption due to full disk errors, preventing conflicting mappings, and fixing NullPointerExceptions during indexing. It also resolved problems with children aggregation missing documents, tribe node configuration requirements, and performance regression after index deletion. Version 2.0.2 saw some of these fixes backported, while 1.7.4 focused on improving delayed shard allocation for efficient recovery. Users were encouraged to upgrade to these versions, particularly 2.1.1, to benefit from enhanced performance and stability, and to report any issues via GitHub or social media.
Dec 17, 2015 691 words in the original blog post.
Advanced Energy Economy (AEE) is a coalition of businesses aiming to make energy secure, clean, and affordable by influencing federal and state policies to open markets for various advanced energy technologies. AEE developed PowerSuite, a platform that aggregates and makes searchable over 45 million pages of regulatory documents from Public Utilities Commissions (PUCs) across the United States, to assist stakeholders in navigating complex state-level regulations that impact $100 billion in annual investments. Initially using PostgreSQL for data management, AEE transitioned to Elasticsearch with the help of Found, which provided a scalable and efficient search solution tailored to their expanding data needs. This shift allowed AEE's small team to focus on delivering impactful user features and advanced analytics without the burden of managing infrastructure, thus enhancing their ability to transform the energy policy landscape. The organization's leadership includes Eric Fitz, an experienced engineer and product developer, and Charlie Forcey, a seasoned web developer and advocate for advanced energy, both of whom bring a wealth of knowledge and expertise to AEE's initiatives.
Dec 17, 2015 1,404 words in the original blog post.
Martin Smith, a DevOps Engineer at Rackspace, provides a detailed guide on deploying Elasticsearch 2.0.0 with Chef, a configuration tool used to manage server settings and automate system configuration. The guide discusses the integration of the new refactored Elasticsearch cookbook for Chef, which simplifies moving from a demo environment to a production-ready cluster. It covers installing and configuring an Elasticsearch node using Chef resources, including setting up a system user, managing package installations, configuring instances, and running Elasticsearch as a system service. Additionally, the article explains how to install Elasticsearch plugins and use Chef's resources to ensure seamless updates and management of server configurations, enabling automation of the setup process. It concludes by verifying the successful deployment of Elasticsearch and mentions future plans to expand on clustering and additional configurations.
Dec 17, 2015 3,740 words in the original blog post.
In the concluding part of a series on building a statistical anomaly detector in Elasticsearch, the author, Zachary Tong, demonstrates how to automate the anomaly detection process using Elasticsearch's Watcher plugin. The initial steps involved creating pipeline aggregations to identify the top 90th percentile of "surprise" values across datasets and using Timelion to visualize these values. The final step integrates these components into the Watcher, enabling real-time alerting and notifications. The process is split into two watches; the first collects surprise data hourly, while the second constructs a dynamic threshold and checks for anomalies, raising alerts if necessary. This approach showcases the power and flexibility of Elasticsearch's tools, allowing for ongoing adjustments and enhancements to the monitoring system.
Dec 16, 2015 2,096 words in the original blog post.
Django Girls, a non-profit organization founded in Berlin in 2014 by Olas Sitarska and Sendecka, aims to empower women through free, one-day programming workshops, focusing on Python and Django. The organization rapidly expanded, with nearly 100 events held in its first year, prompting the need for sustainable growth and the recruitment of Lucie Daeye as the Django Girls Awesomeness Ambassador in September 2015. A key development was receiving sponsorship from Elastic, who raised nearly €15,000, securing Daeye's role and enabling the organization to plan exciting initiatives such as sending swag boxes to workshop organizers and organizing the inaugural Django Girls Summit. These efforts aim to enhance the community's learning atmosphere and provide a platform for organizers to share experiences and strategies. Elastic's support has been instrumental, offering mentorship and sponsorship from the organization's early days, significantly contributing to Django Girls' mission of increasing female representation in the tech community.
Dec 15, 2015 734 words in the original blog post.
Elastic's events and meetups are bustling with activity around the world, drawing to a close with the Elastic{ON} Tour in Tokyo on December 16, 2015. While this marks the end of the current tour, anticipation builds for the Elastic{ON}16 conference in San Francisco, set for February. The week features a variety of meetups across North America, Europe, and Asia, covering topics such as the resiliency of Elasticsearch, its integration with Adobe's Creative Cloud, and the use of Spark and Elasticsearch by eXelate for user analysis. Participants are encouraged to share their experiences on social media and stay engaged with future events, with opportunities to host meetups or present on Elastic's technology stack, including Beats, Elasticsearch, Logstash, and Kibana.
Dec 14, 2015 317 words in the original blog post.
The Logstash project, known for its versatility in data ingestion, has seen significant growth with around 2 million downloads in six months, driven by its extensive ecosystem of over 200 plugins. This success is largely due to community contributions, including issue reporting, suggestions, and new plugin development. To further enhance community involvement and maintain plugin quality, Logstash has introduced Community Maintainers, who will oversee specific plugins and ensure user satisfaction. The first group of seven maintainers, with diverse backgrounds in software engineering and open-source contributions, has already been announced. They are responsible for various plugins and are open to community interaction for questions and feedback. The initiative encourages others to join as maintainers, offering support from the Logstash team and opportunities for greater involvement in open-source projects.
Dec 14, 2015 1,228 words in the original blog post.
Dimitrios Liappis addresses the challenge of resizing the root EBS volume on Linux-based Amazon AMIs, particularly for Debian instances, where increasing the default EBS size does not automatically expand the root filesystem. This issue arises because Debian AMIs lack certain tools like growpart that facilitate automatic resizing, unlike CentOS AMIs which include the necessary packages. To resolve this, Liappis suggests using the 'parted' tool with an init script to resize the root partition non-interactively upon launching a new instance. This script detects discrepancies between the EBS volume and partition sizes, resizes the partition using a specific command, and then enables the filesystem to expand automatically on reboot with cloud-init tools. He also notes that other AMIs, such as CentOS6 and SLES11.x, face similar issues and recommends using growpart where available. The solution involves creating a custom AMI that includes an init script to automate these steps, ensuring the full use of the allocated EBS volume.
Dec 14, 2015 1,217 words in the original blog post.
Nuxeo, a content repository provider, has integrated Elasticsearch to enhance the search capabilities and performance of its platform. Initially, Nuxeo faced challenges with Apache Lucene and a hybrid system involving Python/Zope, which led to limitations in handling complex queries and scalability issues. Transitioning to a Java-based platform, Nuxeo adopted Elasticsearch to address these limitations, offering a hybrid storage solution where SQL handles the primary index while Elasticsearch provides an asynchronous, full-featured index. This integration allows queries to be executed efficiently, depending on the configuration, and significantly improves performance, as demonstrated by benchmarks showing increased throughput and reduced response times. Elasticsearch's integration has also introduced additional features like faceted search, advanced indexing, and an audit trail, enabling real-time data analytics and configurable dashboards. Future plans include upgrading to Elasticsearch 2.0, enhancing security features, and leveraging new Elasticsearch functionalities to further improve the platform's capabilities.
Dec 14, 2015 1,923 words in the original blog post.
Elastic{ON} Tour London 2015 was a sell-out event that gathered Elasticsearch, Logstash, and Kibana experts to explore various applications of these technologies across different industries. Jay Chin of Excelian presented a successful grid project for a major investment bank that utilized Elasticsearch for analytics on grid performance, demonstrating the effective collaboration between Excelian and Elastic engineers. The conference featured discussions on the Elastic roadmap, new features of Elasticsearch 2.0, and real-world applications, including media analytics at The Guardian and a search engine at Goldman Sachs. Notable mentions included NASA's use of Elasticsearch for Mars Rover data analysis and the Victoria and Albert Museum's plan to analyze visitor data. The event provided attendees with insights into Elasticsearch's wide-ranging use cases and facilitated networking opportunities, including discussions with Elastic's creator Shay Banon and co-founder Uri Boness on trends and challenges in the financial services industry.
Dec 09, 2015 852 words in the original blog post.
The text discusses the considerations and challenges involved in deciding whether to store new data in a new type within an existing index or in a new index altogether in Elasticsearch. It highlights the inefficiencies of overusing types, drawing a comparison with relational databases which previously led to misunderstandings. An index in Elasticsearch is stored in shards, and managing a large number of small indices can be inefficient due to the fixed overhead associated with Lucene indices. Types help reduce the number of indices by allowing different data types to be stored within the same index, but they come with limitations such as the need for consistent field configurations across types and issues with sparsity in Lucene indices. The decision to use indices or types depends on factors like data mapping similarity, document volume, and hardware capabilities, with a recommendation to limit shard numbers for optimal resource management. The article notes that there are fewer use cases for multiple types within the same index than might be expected and advises careful consideration of index and shard configuration to maintain efficiency.
Dec 09, 2015 993 words in the original blog post.
Mytaxi, Europe's leading taxi app, leverages the Elastic Stack to maintain app responsiveness and manage extensive logs, supporting a backend of approximately 50 microservices. Initially using a simple Elastic setup, exponential growth necessitated a new logging cluster by mid-2015, addressing increased data storage and performance demands. The new architecture, implemented with AWS and Ansible, features ten nodes of m1.xlarge instances, enabling efficient log storage and retrieval for up to 90 days. This setup supports seamless service migrations, as demonstrated during a significant backend shift to Docker containers on AWS ECS, monitored closely using Kibana for real-time service performance. The Elastic Stack not only provides a comprehensive system overview but also empowers developers to analyze bugs and assess changes' impacts, aligning with the company's goal to enhance the taxi experience across Europe.
Dec 09, 2015 830 words in the original blog post.
In a blog post by Fabian Hueske, the process of building real-time dashboard applications using Apache Flink, Elasticsearch, and Kibana is detailed, highlighting the architecture and implementation of a stream data analytics solution. The architecture leverages Apache Flink for stream processing, Elasticsearch for data storage with low latency, and Kibana for data visualization. The post explains the features of Apache Flink, such as its support for event time and out-of-order streams, expressive APIs, and fault tolerance, which make it suitable for streaming applications. A demo application is described, which analyzes taxi ride events in New York City, showcasing how Flink's DataStream API can be used to compute passenger counts at various locations using sliding window operations. The integration with Elasticsearch and Kibana allows for real-time data visualization, and the post provides steps for setting up and configuring these components. Hueske concludes by encouraging readers to experiment with the demo and explore the capabilities of the tools used.
Dec 07, 2015 2,648 words in the original blog post.
Elastic's global events schedule for December 2015 highlights several key gatherings, including the Elastic{ON} Tour stops in Sydney and Melbourne, marking the tour's first visit to this region with free events featuring product overviews, use case deep dives, and workshops. In North America, the Elastic Denver User Group's meetup will focus on Elasticsearch 2.0, and a series of meetups are set across Europe, including events in Paris, Warsaw, Munich, Karlsruhe, Kyiv, and Rome. Elastic invites participation and offers support for those interested in hosting meetups or giving talks related to its products.
Dec 07, 2015 283 words in the original blog post.
Atlas, a statistical anomaly detection system developed for Elasticsearch, leverages pipeline aggregations to distill large datasets into key metrics, focusing on the 90th percentile of "surprise" values, or deviations from the moving average, to identify anomalies over time. The system employs TimeLion for post-processing, graphing these variations, and setting alerts when deviations exceed three standard deviations above the moving average. This setup enables efficient anomaly monitoring without examining vast quantities of data directly, by highlighting significant variance changes that suggest disruptions. Despite the limitations of current pipeline aggregations, such as their inability to select the "last" surprise value, Atlas remains effective, even with skewed data distributions, by relying on its ability to track and respond to significant shifts in data variance. Future enhancements may include integrating alerting systems like Watcher to automate notifications for detected anomalies.
Dec 02, 2015 1,734 words in the original blog post.
Organizations often need to manage and replicate data across multiple regions due to local security, privacy, and performance requirements, which presents challenges such as high availability, fault tolerance, and varying network conditions. The blog discusses potential architectures for a creative agency with offices in New York and London, where media assets are created locally and need to be accessible across data centers. A single shared Elasticsearch cluster is discouraged due to potential synchronization issues caused by network failures. Independent Elasticsearch clusters with Tribe Nodes or a shared Kafka cluster offer alternatives but come with their own drawbacks, such as search latency and network dependency. The recommended architecture involves maintaining independent Elasticsearch and Kafka clusters in each data center, with data synchronization achieved through Kafka MirrorMaker or Logstash. This approach allows for reliable data replication, ensures independent operation in the event of network failures, and supports disaster recovery by allowing each location to serve as a backup for the other. While Elasticsearch does not inherently support consistent replication over high latency networks, combining it with Kafka provides a robust solution for the use case.
Dec 01, 2015 1,090 words in the original blog post.