Home / Companies / Voxel51 / Blog / June 2024

June 2024 Summaries

6 posts from Voxel51

Filter
Month: Year:
Post Summaries Back to Blog
The June '24 AI, Machine Learning, and Computer Vision Meetup covered topics such as leveraging pre-trained text2image diffusion models for zero-shot video editing, improved visual grounding through self-consistent explanations, and combining Hugging Face Transformer models with image data using FiftyOne. Bariscan Kurtkaya discussed the potential of using pre-trained text-to-image diffusion models for video editing without fine-tuning. Dr. Paola Cascante-Bonilla presented her work on enhancing vision-and-language models' ability to localize objects in images by fine-tuning them for self-consistent visual explanations. Jacob Marks demonstrated how the seamless integration between Hugging Face and FiftyOne simplifies connecting datasets and models, enabling more effective data-model co-development. The next Meetup is scheduled for July 3rd, featuring talks on performance optimization for multimodal LLMs, five handy ways to use embeddings in AI, and responsible and unbiased genAI for computer vision.
Jun 27, 2024 1,291 words in the original blog post.
The CVPR conference highlighted several insightful papers this year. CoDeF tackles the issue of inconsistency in video-to-video translation by representing videos with a flattened canonical image and a deformation field, enabling unprecedented cross-frame consistency. Depth Anything revolutionizes depth estimation using a Dense Prediction Transformer (DPT) architecture, offering unparalleled generality and robustness for zero-shot depth estimation. YOLO-World bridges the gap between real-time closed-vocabulary detection and open-vocabulary object detection by combining a YOLO backbone with semantic information from a CLIP text encoder. DeepCache accelerates diffusion model inference by up to 10x, leveraging consistent high-level features throughout the denoising process. PhysGaussian integrates physical concepts like stress and elasticity into machine learning models for real-time motion synthesis.
Jun 21, 2024 856 words in the original blog post.
In this blog post, Robert Wright shares his experience of building a dataset for an in-cabin monitoring use case using Voxel51's open source AI tool, FiftyOne. He collected data using his wife's iPhone and a car mount, then loaded the data into both the open source and enterprise versions of FiftyOne. By applying various models such as MediaPipe Face Detection and Face Landmarker pipelines, he gained insights into driver distractions, eye gaze, and facial expressions. The author also highlights the collaboration features of FiftyOne Teams, which allowed him to work with Voxel51 ML engineer Allen Lee on this project.
Jun 17, 2024 2,707 words in the original blog post.
The text discusses five interesting papers from CVPR 2024. CoDeF is a technique that overcomes the challenge of breaks in temporal consistency in video editing/translation by representing any video with a flattened canonical image and a deformation field. Depth Anything revolutionizes depth estimation using just a single image, offering unparalleled generality and robustness for zero-shot depth estimation. YOLO-World bridges the gap between real-time closed-vocabulary detection and open-vocabulary object detection by introducing semantic information via a CLIP text encoder. DeepCache accelerates diffusion model inference by up to 10x with minimal quality drop-off, leveraging high-level feature consistency throughout the denoising process. PhysGaussian is a physics-based machine learning approach that embeds physical concepts like stress, plasticity, and elasticity into the model itself for simulating dynamics.
Jun 14, 2024 991 words in the original blog post.
CVPR 2024 is happening next week with Voxel51 team members attending. The event will feature interviews with authors presenting research in vision-language models, 3D computer vision, and diffusion models. Visitors can learn about the open source FiftyOne project at booth #1519, which helps improve data quality and boost model performance. There are also new releases of FiftyOne since last year's CVPR. Voxel51 is recruiting for various roles and will be hosting two workshops with Chief Scientist Jason Corso.
Jun 10, 2024 987 words in the original blog post.
The May '24 AI, Machine Learning, and Data Science Meetup featured presentations on fine-tuning Llama2 for autonomous agents, integrating Hugging Face Transformer models with image data using FiftyOne, multi-modal visual question answering (VQA) using UForm tiny models with Milvus vector database, strategies for enhancing the adoption of open source libraries, and more. The event also included a charity donation to Heart to Heart International based on attendee votes. Upcoming Meetups are scheduled for June 27th, featuring topics such as leveraging pre-trained text2image diffusion models for zero-shot video editing, improved visual grounding through self-consistent explanations, and more.
Jun 03, 2024 1,226 words in the original blog post.