Home / Companies / Voxel51 / Blog / June 2026

June 2026 Summaries

20 posts from Voxel51

Filter
Month: Year:
Post Summaries Back to Blog
OlmoEarth v1.2, developed by Allen AI, is a foundation model for satellite imagery that significantly reduces computational costs without sacrificing accuracy, achieving 3× fewer GPU-hours during training and 2.9× fewer operations during inference compared to its predecessor, v1. This efficiency is achieved through advanced engineering techniques such as token collapsing and improved masking. It supports various satellite data including Sentinel-1, Sentinel-2, and Landsat, and can operate on Apple Silicon without a GPU. FiftyOne, Voxel51's open-source toolkit, complements OlmoEarth by visualizing and managing the model's high-dimensional embeddings, allowing users to explore satellite imagery through UMAP-reduced vectors. This integration enables unsupervised clustering and provides a synchronized view of embedding space, image grids, and geospatial maps. Both OlmoEarth and FiftyOne are open-source, with OlmoEarth's models and dataset openly available and including a license that prohibits military and extractive-industry applications.
Jun 30, 2026 1,288 words in the original blog post.
FiftyOne Annotation, developed by Voxel51, is a comprehensive platform designed to enhance data annotation processes for AI systems, particularly in the context of physical AI teams handling complex and diverse datasets. It integrates data curation, annotation, and model evaluation into a single system, aimed at producing high-quality labels efficiently and cost-effectively. The platform offers configurable workflows for 2D, 3D, and video data, enabling annotators to apply various label types and ontologies consistently across projects, while role-based task assignments streamline project management. Key features include Agentic Labeling, which uses Visual Language Models (VLMs) to automate initial labeling, allowing human annotators to focus on refining these outputs. Smart Data Selection techniques prioritize impactful samples for labeling, reducing wasted resources. Intelligent Review and QA mechanisms identify errors before they reach model training, ensuring data quality. By closing the loop with model evaluation, FiftyOne helps teams iteratively improve data selection and annotation processes, ultimately accelerating model development and enhancing performance.
Jun 30, 2026 2,305 words in the original blog post.
The article explores the development of a plugin within FiftyOne, an open-source tool for creating high-quality datasets and evaluating computer vision models, specifically designed for visualizing complex autonomous driving data. Inspired by nuReasoning's long-tail driving clips, the plugin is crafted to synchronize and display multiple data modalities—such as camera frames, LiDAR views, and live maps—captured simultaneously by various sensors on a self-driving car. This synchronization allows for real-time debugging and analysis of multimodal data, enabling users to better understand spatial relationships and sensor fusion issues. The plugin leverages FiftyOne's flexible plugin framework, which supports custom logic and visualization, and demonstrates how grouped datasets can be effectively utilized to provide a cohesive view of sensor data, enhancing the debugging and evaluation process for autonomous driving and other applications like robotics and aerial mapping.
Jun 29, 2026 2,382 words in the original blog post.
SAM 3, short for Segment Anything Model 3, is a Meta foundation model introduced in November 2025 that revolutionizes data annotation by transforming text prompts into comprehensive segmentation across images and video. Unlike its predecessors, SAM and SAM 2, which required manual clicks to segment objects, SAM 3 utilizes open-vocabulary concept prompting, allowing it to identify and segment every instance of a described concept, such as "yellow school bus," in one go. This advancement significantly reduces the time and cost associated with manual segmentation, making the annotation process more efficient. However, SAM 3 does not eliminate the need for human involvement, as it cannot determine which data should be labeled or ensure the accuracy of the labels, especially in complex or specialized datasets. The model shifts the annotation bottleneck to the stages of review and selection, requiring human judgment to verify and refine the automated results. SAM 3.1, released in March 2026, further enhances the model's capabilities by introducing Object Multiplex for faster real-time video segmentation, but the core function of promptable concept segmentation remains unchanged. Despite its advancements, SAM 3 is positioned as a tool to accelerate and streamline the annotation process rather than replace human annotators entirely, emphasizing the ongoing need for human oversight in ensuring quality and relevance in data labeling.
Jun 29, 2026 2,162 words in the original blog post.
Physical AI refers to artificial intelligence systems that interact with the real world through sensors and actuators, used in technologies such as autonomous vehicles, robots, and drones. Unlike digital AI, which operates in the virtual realm, physical AI must accurately navigate a constantly changing and unpredictable environment, where mistakes have tangible consequences. The primary challenge lies in handling multimodal data from various sensors like cameras, lidar, and radar, which must be synchronized, curated, and evaluated to improve system performance. A significant insight from Voxel51's 2026 State of Visual and Physical AI report is that data, particularly rare, safety-critical events, is more crucial to success than model size. Physical AI systems operate on a "curate, annotate, and evaluate" loop, requiring proprietary data to maintain a competitive edge. The need for comprehensive evaluation processes underscores the importance of data as the central problem in deploying reliable physical AI solutions.
Jun 28, 2026 3,090 words in the original blog post.
nuReasoning is an innovative autonomous driving dataset developed by Motional in collaboration with UCLA, designed to address complex, rare driving scenarios by incorporating reasoning-based annotations. The dataset, consisting of approximately 20,000 clips from diverse U.S. locations, provides multi-modal data including camera views, LiDAR, and maps, with annotations for spatial, decision, and counterfactual reasoning. It aims to improve the understanding of autonomous vehicle decision-making by showing not only what decisions were made but also why and what alternatives were considered. FiftyOne, an open-source platform from Voxel51, enhances the dataset's utility by enabling interactive exploration of these clips, allowing users to visualize and understand the reasoning processes behind each driving decision. This integration provides a detailed, visual insight into machine decision-making, making it a valuable resource for researchers and developers focused on autonomous systems.
Jun 26, 2026 1,603 words in the original blog post.
The article explores the importance of conducting ground truth audits on datasets, using the Pyro-SDIS wildfire smoke detection dataset as a case study. It highlights that every object detection benchmark relies on the assumption that human-created ground truth annotations are correct, an assumption often untested but crucial for model quality. The audit, performed using FiftyOne, a dataset curation and evaluation platform, revealed that while the Pyro-SDIS dataset is geometrically clean with no severe train/validation leakage, it suffers from fixed-camera redundancy. The study emphasizes the need for deduplication and camera-aware splits to ensure accurate evaluations rather than splitting frames randomly. It also suggests prioritizing human review on annotation errors supported by multiple independent methods, as single-method findings may be misleading. The article underscores the necessity of triangulating findings across different methods and models to ensure credibility in dataset audits, advocating for a non-destructive, reproducible audit process that adapts to domain-specific challenges like differentiating smoke from fog or clouds.
Jun 26, 2026 3,885 words in the original blog post.
Medical data annotation, particularly in the context of medical imaging, often involves expert disagreement, which is traditionally seen as noise to be averaged out using algorithms like STAPLE. However, in the era of foundation models, such as UNI2 and MedSAM2, where datasets are smaller and more specific, this disagreement should be viewed as a valuable signal rather than a problem. Treating disagreement as a first-class signal can enhance model reliability by identifying edge cases and potential failures. This approach requires explicit representation of disagreements, exploring them through embeddings, and careful curation of datasets to maintain high-quality annotations. Furthermore, regulations like the EU AI Act and FDA frameworks demand comprehensive documentation of annotation quality, making it crucial for teams to adopt workflows that preserve individual annotations and disagreement data. By maintaining detailed records and focusing on disagreement, teams can ensure compliance and improve the performance and reliability of AI models in healthcare settings.
Jun 26, 2026 1,635 words in the original blog post.
Collecting safety-critical off-road data poses challenges due to the difficulty and danger involved in capturing adverse weather conditions and obstacles in real-world scenarios. The integration of ComfyUI with FiftyOne offers a solution by generating synthetic data to fill these gaps, thereby allowing for the creation of diverse and representative datasets without risking real vehicles. This process involves identifying data gaps within FiftyOne, using ComfyUI to generate synthetic conditions and obstacles, and curating the produced data to ensure its relevance and accuracy. The use of the STONE dataset exemplifies this method, where synthetic generation complements real-world data collection without replacing it, as each synthetic frame is traceable back to its real-world origin. This approach, while maintaining the integrity of the original 3D voxel data, enhances model training by enabling the inclusion of otherwise inaccessible conditions and obstacles, effectively turning ComfyUI and FiftyOne into a cohesive data generation and curation loop.
Jun 25, 2026 3,624 words in the original blog post.
The article introduces the Annotation Styles plugin for FiftyOne, an open-source platform for dataset curation and model evaluation in computer vision. This plugin allows users to dynamically alter the visual representation of annotation styles directly within FiftyOne's dataset viewer, enhancing the exploration of data without modifying the dataset or saving new files. Unlike Roboflow's supervision library, which renders styles into new pixel images, the Annotation Styles plugin applies styles live, offering immediate updates and interactivity across thousands of samples without re-rendering. Users can select from various styles for different label types, including detections, keypoints, polylines, and segmentation masks, facilitating model debugging, privacy review, and dataset presentation. The plugin supports side-by-side comparisons with different styles applied to different fields simultaneously, making it a powerful tool for improving the understanding and presentation of computer vision data.
Jun 24, 2026 1,791 words in the original blog post.
Data annotation, a critical process in machine learning, involves attaching structured labels to raw data, enabling models to learn the patterns these labels describe. By curating the most relevant data before labeling, teams can ensure efficient use of resources, as labeling redundant data wastes budget and doesn't enhance model performance. Various annotation types, such as classification, bounding boxes, and segmentation, cater to different model learning needs. Modern annotation workflows incorporate automated and agentic labeling techniques, where AI models assist in preliminary labeling, requiring humans to focus on refining and correcting outputs. As machine learning architectures and compute become commoditized, the quality and strategic choice of labeled data become pivotal in determining a project's success. Reports highlight that most failures in AI projects stem from undervalued data quality, and exceptional teams prioritize data work and maintain a continuous feedback loop in their workflows to enhance model accuracy effectively.
Jun 23, 2026 2,809 words in the original blog post.
In the pursuit of reducing costs, many annotation managers prioritize high throughput in data labeling, often at the expense of data quality, leading to underperforming machine learning models. This approach, akin to an accounting error in baseball history, can result in "phantom hits" where incorrect or duplicated data inflates perceived progress without genuine improvement. The article argues that focusing solely on cost-per-label metrics masks the degradation of model performance, emphasizing that annotation is fundamentally a quality issue. Instead of simply increasing labeled data, which may only reinforce existing knowledge, the text advocates for strategic data curation that targets rare, critical samples to significantly enhance model accuracy. Research demonstrates that curated data improves models more effectively than raw data volume increases, and the cultural undervaluation of data work results in compounding failures. To ensure robust machine learning models, teams should measure model improvement per dollar and prioritize accurate, high-quality data labeling over sheer volume.
Jun 23, 2026 1,776 words in the original blog post.
Exploring the PointMotionBench Benchmark in FiftyOne highlights the capabilities of the MolmoMotion model from Ai2, which distinguishes itself by forecasting 3D motion from a single frame and plain-language instructions, unlike traditional models that only narrate past movements. The PointMotionBench, integral to this exploration, comprises .npz track files and JSON captions from datasets like DAVIS, HOT3D, and WorldTrack, serving as a critical resource for training and scoring object-centric 3D motion forecasting. By using FiftyOne to transform these raw tracks into a browsable dataset, users can gain a deeper understanding of the benchmark's contents, challenges, and data diversity. This process includes visualizing ground truth and predicted tracks, which helps users grasp the complexities of 3D motion prediction and evaluate model performance effectively. The article suggests utilizing a GPU to run MolmoMotion offline for predictions, enhancing the demo experience by comparing predicted and ground-truth motion side by side, and sorting clips by error to identify model weaknesses. The notebook initially focuses on the DAVIS split, with potential expansion to HOT3D and WorldTrack once access is obtained, thereby offering a comprehensive toolset for understanding and improving 3D motion forecasting models.
Jun 22, 2026 1,350 words in the original blog post.
TruckDrive is a long-range highway driving dataset designed to reveal the limitations of current autonomous driving models, particularly their inability to accurately perceive and respond to objects beyond 150 meters, a critical range for highway speed stopping. Unlike city datasets with limited perception ranges, TruckDrive, developed by Torc Robotics and Princeton, extends perception up to 400 meters using a sophisticated sensor suite, including long-range LiDARs and cameras. When evaluated with state-of-the-art models, performance drastically declines beyond 150 meters, highlighting a significant "long-range gap" in current autonomous systems. The dataset is integrated into FiftyOne, a visualization tool that allows users to explore these limitations interactively by tagging objects with their distance from the ego vehicle, thereby transforming abstract performance metrics into tangible insights. The dataset's visualization in FiftyOne emphasizes the importance of long-range perception, with features that allow users to toggle between 3D point clouds and camera views to better understand the depth and spatial relationships. Limitations in visualizing radial velocity due to sensor-specific frames are noted, but future work could address these through enhanced data integration and visualization techniques. The TruckDrive dataset is released under a non-commercial license, inviting further exploration and development to address the long-range perception gap in autonomous driving.
Jun 17, 2026 1,360 words in the original blog post.
A Reddit post about a lost turbine RC plane in the desert led to the use of Voxel51's FiftyOne tool, showcasing its powerful plugin system for handling complex computer vision challenges. The plane's owner had 3,000 high-resolution images to analyze, and instead of using a basic script, the workflow involved transforming these images into a queryable dataset, emphasizing the dataset as the core asset for analysis. FiftyOne's platform, including its Brain for embeddings and a unique plugin system, enabled the conversion of raw data into ranked shortlists, making it easier to identify potential matches. The plugin system allowed the creation of reusable and shareable tools that could operate both within a coding environment and as an intuitive user interface, demonstrating the efficiency and flexibility of the platform. Ultimately, the approach significantly reduced the manual review time, highlighting the advantages of treating datasets as dynamic, interactive components rather than static outputs.
Jun 17, 2026 1,636 words in the original blog post.
The VAND 4.0 Kaputt Challenge, held at CVPR 2026 in Honolulu, highlighted the complexities of defect detection in real-world retail logistics through Amazon's expansive Kaputt dataset. This dataset, introduced at ICCV 2025, comprises over 230,000 images and 29,000 defective instances, capturing the unpredictable nature of retail environments with varying object poses and appearances. The challenge encouraged teams to bridge the gap between controlled benchmarks and practical applications, with submissions demonstrating progress in anomaly detection amidst real-world conditions. The FiftyOne library facilitated the exploration of this dataset, offering tools for interactive filtering, similarity search, and visualization to enhance understanding and development of robust computer vision models. Sponsors, including Amazon, Voxel51, Intel, and MVTec, contributed prizes and resources, underscoring the significance of efficiency and real-world applicability in advancing the field.
Jun 16, 2026 1,737 words in the original blog post.
D4RT (Dynamic 4D Reconstruction and Tracking) is a groundbreaking model developed by Google DeepMind, University College London, and the University of Oxford, which won the best paper award at CVPR 2026 for its innovative approach to 4D scene reconstruction. Unlike traditional pipelines that rely on multiple specialized models for depth, optical flow, and camera pose, D4RT utilizes a single query interface to efficiently encode video into a latent scene representation, enabling dynamic and static object tracking without separate fusion steps. The model sets a new state of the art across various 4D reconstruction benchmarks, particularly excelling in tracking moving objects where previous methods like VGGT struggle. Although the model weights have not yet been released, a companion notebook using the FiftyOne toolkit simulates D4RT outputs, providing an interactive way to explore its capabilities and visually demonstrate the model's unified approach in handling dynamic scenes.
Jun 11, 2026 3,211 words in the original blog post.
FiftyOne Enterprise 2.19.0 introduces significant enhancements, including the FiftyOne Agent, which enables users to interact with visual datasets using natural language, supported by over 100 LLM providers. Enhancements in in-app annotation include versioned ontologies and AI-assisted masks, allowing for more efficient and dynamic labeling processes. A new Observability feature provides comprehensive visibility into delegated runs with live metrics, logs, and telemetry. The release also includes quality-of-life improvements such as the addition of Segment Anything 3 to the Model Zoo and more flexible panel options, alongside security upgrades addressing CVEs and dependency vulnerabilities. These upgrades unify data, models, and teams in a secure platform designed to enhance model performance and user experience.
Jun 09, 2026 1,330 words in the original blog post.
Autonomous driving research faces challenges due to fragmented datasets, each with its own formats and conventions, making cross-dataset training cumbersome. The open-source library py123d addresses this by converting data from various autonomous vehicle datasets into a unified Apache Arrow format, allowing seamless access to cameras, lidar, HD maps, and labels with a single API. It utilizes efficient memory-mapped reads and supports multiple sensor codecs without duplicating storage. Further enhancing this process, the FiftyOne library provides an interactive platform for exploring and comparing datasets, offering features such as label filtering, dataset-scale querying, and cross-dataset comparisons, all within a browser interface. Together, py123d and FiftyOne streamline the ingestion and understanding of autonomous driving data, allowing researchers to efficiently manage, visualize, and analyze large datasets across different sources.
Jun 08, 2026 1,326 words in the original blog post.
MR-RATE is a substantial vision-language dataset comprising over 700,000 brain and spine MRI volumes, paired with radiology reports and structured metadata, designed to advance research in medical imaging. Made available on HuggingFace, this dataset supports a wide range of clinical and research applications, enabling exploration and analysis using FiftyOne, an open-source visual dataset curation tool. FiftyOne facilitates interactive exploration of this massive dataset, allowing users to filter images by metadata, view detailed radiology reports, and conduct visual similarity searches. The tutorial accompanying the dataset demonstrates how to import MR-RATE into FiftyOne, convert 3D MRI volumes into 2D images, and use a pretrained ResNet18 model to generate embeddings for visualization and nearest-neighbor search. This process makes the dataset accessible and useful for developing diagnostic tools, training vision-language models, and studying neurological conditions, offering a low barrier to entry for significant clinical research.
Jun 03, 2026 1,595 words in the original blog post.