July 2025 Summaries
20 posts from Google Cloud
Filter
Month:
Year:
Post Summaries
Back to Blog
The launch of Veo 3 Fast presents an optimized model for speed and cost-effectiveness, allowing developers to efficiently create high-quality video outputs. Veo 3 Fast is tailored for business use cases such as programmatic advertising, rapid prototyping, and large-scale content creation, offering both text-to-video and image-to-video capabilities at $0.40 per second with audio. The newly introduced image-to-video feature allows for the creation of dynamic video sequences from a single still image, maintaining consistency and providing creative flexibility through precise prompting and seamless API integration. This feature, available via the Gemini API, is priced at $0.75 per second with audio and is designed to enhance video editing experiences by generating cinematic-quality videos with minimal effort.
Jul 31, 2025
547 words in the original blog post.
The general availability of the Gemini Embedding text model has spurred rapid adoption among developers for creating advanced AI applications, expanding beyond traditional uses such as classification and semantic search to include context engineering for providing AI agents with complete operational contexts. This model's embeddings effectively integrate vital information into a model's working memory, enhancing capabilities across various industries. For instance, Box utilizes it to extract insights from complex multilingual documents, achieving a recall increase of 3.6%. Financial technology company re:cap reports improved classification accuracy of B2B transactions, with a 1.9% increase in F1 score, while Everlaw benefits from precise semantic matching in legal discovery, outperforming other models with 87% accuracy. Additionally, Roo Code enhances codebase searches, and Mindlid's AI wellness companion delivers personalized support with improved relevance and speed. Interaction Co.'s AI email assistant, Poke, uses it for efficient email context retrieval, achieving a 90.4% reduction in embedding time. These diverse applications demonstrate the model's potential in driving significant performance gains and efficiency in AI systems.
Jul 30, 2025
695 words in the original blog post.
LangExtract is an open-source Python library designed to efficiently extract structured information from unstructured text using large language models (LLMs), providing developers with a powerful tool for information extraction across various domains such as medicine, finance, and law. It emphasizes the traceability of extracted data by mapping entities back to their source text and supports interactive visualization to facilitate evaluation and verification. The library allows users to define their desired outputs through custom instructions and few-shot examples, ensuring reliable and consistent data structuring. By leveraging controlled generation and optimized information extraction techniques, LangExtract can handle complex documents through chunking, parallel processing, and context-specific extraction. It supports various LLM backends, including cloud-based and on-device models, and can incorporate the inherent world knowledge of LLMs to enhance extracted information. LangExtract is demonstrated through examples like medication extraction from clinical text and structured radiology reporting, showcasing its potential to improve data clarity and interoperability in specialized fields.
Jul 30, 2025
1,149 words in the original blog post.
JAX is increasingly being utilized by developers across various computational fields, extending its application beyond large-scale AI to include domains such as robotics, where it enhances simulation, control, and learning-based methods. Max Muchen Sun, a Robotics Ph.D. candidate at Northwestern University, exemplifies how JAX addresses complex challenges in robotics, particularly with computational efficiency in control algorithms and integrating model-based and learning-based approaches. Sun's transition from traditional tools to JAX features like vmap and scan highlights the framework's ability to facilitate parallelization and accelerate trajectory simulations. His work demonstrates JAX's strengths in merging model-based and learning-based pipelines, as seen in projects involving flow matching and multi-agent cooperation, implemented using JAX-native tools. Sun developed the LQRax package, showcasing JAX's capability to support GPU acceleration and differentiable LQR, emphasizing its role in real-time control and complex planning. The JAX ecosystem, supported by tools such as Brax, MJX, and JaxSim, continues to grow, offering robust solutions for robotics and trajectory optimization, indicating a promising future for JAX in advancing intelligent robotic systems.
Jul 29, 2025
1,140 words in the original blog post.
As enterprises strive to operationalize AI, integrating large language models (LLMs) into existing API ecosystems while ensuring security, governance, and compliance is a significant challenge. Apigee, Google Cloud's API management platform, plays a crucial role in this integration process by enhancing the security, scalability, and governance of gen AI agents within applications. The Model Context Protocol (MCP) has become a prominent method for integrating discrete APIs, yet its rapid evolution means it doesn't fully address enterprise needs for authentication, authorization, and observability. Apigee provides an open-source example of an MCP server with robust API security features, demonstrating how enterprises can leverage these tools to secure, scale, and govern their AI interactions. This setup bridges the gap between managed APIs and exploratory AI interactions, making it adaptable to changes in the MCP standard. Apigee offers a GitHub repository with resources to guide users in deploying the reference MCP Serving architecture, emphasizing its commitment to evolving alongside the AI landscape and supporting enterprises in their AI journeys.
Jul 24, 2025
656 words in the original blog post.
Season 5 of Google's People of AI podcast shifts its focus to the builders of Artificial Intelligence, delving into the rapidly evolving role of developers, startups, and industry leaders in shaping AI's future. This season, hosted by Ashley Oldacre and veteran podcaster Christina Warren, offers insights into the groundbreaking advancements in AI technologies such as convolutional neural networks, transformers, and large language models (LLMs). The podcast features interviews with key figures like Clement Farabet from Google DeepMind, discussing the potential and challenges of developing AI agents that are both powerful and responsible. By exploring the stories and innovations of AI builders, listeners gain a deeper understanding of how AI is transforming industries and daily life, encouraging creative thinking and problem-solving in a world increasingly influenced by AI.
Jul 24, 2025
424 words in the original blog post.
Opal, an experimental tool from Google Labs, enables users to create and share AI mini apps by chaining together prompts, models, and tools using natural language and visual editing. Designed to simplify the process of building AI applications, Opal allows for the creation of workflows without requiring any coding knowledge, making it ideal for prototyping AI ideas, demonstrating proofs of concept, and enhancing productivity. Now in a US-only public beta, Opal encourages community involvement from the start and offers features such as a visual editor for easy modifications and a demo gallery with starter templates. Once an app is built, it can be shared with others for immediate use via their Google accounts, supporting a new way of creating with AI to empower creators and innovators.
Jul 24, 2025
402 words in the original blog post.
Firebase Studio, a cloud-based AI workspace, is enhancing agentic AI development with new updates unveiled at I/O Connect India, aimed at facilitating the creation of AI-powered apps using popular frameworks like Flutter, Angular, React, and Next.js. The updates feature AI-optimized templates that leverage the Gemini CLI for faster development, streamlined integration with Firebase backend services, and enhanced flexibility through workspace forking. These improvements allow developers to build, experiment, and collaborate more efficiently by offering autonomous Agent modes, greater control over evolving codebases, and improved prompts for refining project ideas. The platform's new features, such as increased project upload sizes and enhanced AI-assisted development tools, aim to empower both seasoned and novice developers to innovate and share their creations globally.
Jul 23, 2025
854 words in the original blog post.
Gemini 2.5 Flash-Lite, the latest release in the Gemini 2.5 model family, combines high speed and cost-efficiency, priced at $0.10 per million input tokens and $0.40 per million output tokens, making it ideal for high-volume tasks like translation and classification. It offers a significant reduction in latency and power consumption compared to previous models, with improved performance across various benchmarks such as coding, math, and multimodal understanding. Users benefit from a 1 million-token context window and support for native tools like Google Search and Code Execution. Notable applications include Satlyt’s space computing platform, HeyGen’s multilingual video content automation, DocsHound’s documentation processing, and Evertune’s brand representation analysis. The model is now available for deployment in Google AI Studio and Vertex AI, with a transition from its preview version scheduled for August 25th.
Jul 22, 2025
516 words in the original blog post.
AI's ability to visually understand images has advanced significantly, evolving from merely identifying object locations with bounding boxes to employing segmentation models that accurately outline object shapes. The latest progression involves open-vocabulary models capable of segmenting objects using less typical labels without a predefined category list. However, the more complex challenge of conversational image segmentation, akin to referring expression segmentation, requires a nuanced understanding of detailed descriptive phrases. This is where Gemini's advanced visual capabilities excel, allowing for intricate queries that go beyond simple labels, enabling the identification of objects based on various relationships, conditional logic, abstract concepts, in-image text recognition, and multilingual labels. These capabilities facilitate new applications such as creative media editing, safety compliance monitoring, and nuanced insurance damage assessments. For developers, Gemini offers a flexible language approach and a simplified development experience, allowing for the creation of sophisticated vision applications without the need for specialized models. Users can leverage these features in Google AI Studio or a Python environment, supported by comprehensive documentation and a developer forum for community interaction.
Jul 21, 2025
874 words in the original blog post.
Veo 3, introduced at Google I/O 2025 and now available in paid preview via the Gemini API and Vertex AI, is a cutting-edge video generation model designed to produce high-quality videos with synchronized audio. This innovative tool supports text-to-video and image-to-video capabilities, enabling developers to create cinematic narratives and dynamic character animations with realistic physics and rich audio elements. Companies like Cartwheel and Volley are already leveraging Veo 3 to create 3D animations and in-game video cut-scenes, enhancing their creative outputs. The model's applications can range from intricate storytelling to interactive gaming experiences, with pricing set at $0.75 per second for video and audio output. Veo 3 is accessible through Google AI Studio, a platform that provides SDK templates and an interactive Starter App for rapid prototyping, and will soon offer a faster, more cost-effective version called Veo 3 Fast. All videos generated with Veo 3 include a digital SynthID watermark to ensure authenticity and responsible usage.
Jul 17, 2025
917 words in the original blog post.
The introduction of the logprobs feature in the Gemini API on Vertex AI provides developers with a transparent view of the decision-making process of language models by displaying probability scores for chosen tokens and their alternatives. This feature, beyond being a debugging tool, enables the development of smarter, more reliable, and context-aware applications by allowing insight into model reasoning, making it suitable for use cases such as confident classification, dynamic autocomplete, and quantitative retrieval-augmented generation (RAG) evaluation. Logprobs, representing the natural logarithm of a token's probability score, highlight the model's confidence in its choices, with scores closer to zero indicating higher confidence. The blog details how to enable and process logprobs, using a step-by-step guide through an "Intro to Logprobs" notebook, and demonstrates applications like detecting ambiguity in classifications, enhancing auto-complete features with contextual predictions, and evaluating RAG systems by correlating confidence scores with retrieval quality. This innovation presents new opportunities for developers to create more transparent and adaptable AI-driven applications.
Jul 16, 2025
1,229 words in the original blog post.
The Marin project, spearheaded by Stanford's Center for Research on Foundation Models (CRFM), seeks to redefine openness in AI by offering a comprehensive, transparent approach to foundation model development. This initiative includes not just sharing models like Marin-8B-Base and Marin-8B-Instruct but also providing access to code, datasets, methodologies, and training logs under the Apache 2.0 license. The project tackles engineering challenges such as maximizing computational efficiency and ensuring reproducibility using JAX and its components like Levanter and Haliax, which enable scalable and deterministic training processes. Marin's adaptive training journey, dubbed the "Tootsie" process, showcases the robustness of its tools by maintaining bit-for-bit reproducibility despite variations in data and hardware configurations. By promoting transparency and collaboration, the Marin project invites researchers to engage in developing trustworthy AI models, with resources available on platforms like Hugging Face and GitHub, and fostering community interaction via Discord and other channels.
Jul 16, 2025
1,541 words in the original blog post.
Developers highly value the state of "flow," where coding occurs with minimal effort, but this state is often disrupted by various frictions such as documentation review and context-switching. To mitigate these disruptions, updates to the Agent Development Kit (ADK) combined with the Gemini CLI have been introduced, enabling a more seamless and efficient "vibe coding" experience. Central to these improvements is the revamped llms-full.txt file, which serves as a comprehensive and more concise guide to the ADK framework, allowing for better understanding by language models without overwhelming them. This enhanced synergy between ADK and Gemini CLI facilitates the rapid conversion of high-level ideas into functional agents, exemplified by the creation of an AI agent for labeling GitHub issues. Through a streamlined process of ideation, code generation, and iterative improvement, developers can now produce effective agent applications swiftly, maintaining their creative momentum without the usual coding overhead.
Jul 16, 2025
1,278 words in the original blog post.
Gemini Embedding, Google's latest text model known as gemini-embedding-001, is now available through the Gemini API and Vertex AI, offering advanced capabilities for a range of applications including science, legal, finance, and coding. This model has been a leader on the Massive Text Embedding Benchmark (MTEB) Multilingual leaderboard since its experimental release, outperforming previous models and external competitors in tasks like retrieval and classification. It supports over 100 languages and employs the Matryoshka Representation Learning technique for flexible output dimensions, allowing optimization for performance and storage. With both free and paid tiers, developers can access the model through Google AI Studio, and pricing is set at $0.15 per million input tokens. As legacy models will soon be deprecated, users are encouraged to transition to gemini-embedding-001, which promises to unlock new possibilities and will soon support asynchronous processing through the Batch API.
Jul 14, 2025
461 words in the original blog post.
The Apigee team at Google Cloud emphasizes the importance of building robust API ecosystems by providing tools like the Apigee API hub and Developer Portals, which cater to different needs within an organization. The Apigee API hub serves as the central repository or "enterprise truth" for all APIs, focusing on the needs of API producers such as platform engineers and developers, by cataloging APIs and their metadata, ensuring governance, security, and facilitating internal collaboration. Developer Portals, on the other hand, are designed for API consumers, offering a curated, user-friendly interface for discovering and interacting with a selected subset of APIs. These portals rely on the comprehensive data from the Apigee API hub to ensure accurate and effective API discovery, thereby enhancing the consumption experience. Together, the Apigee API hub and Developer Portals form a cohesive system that not only supports API management but also lays the groundwork for AI agent strategies, enabling organizations to unlock the full potential of their APIs and drive innovation.
Jul 14, 2025
862 words in the original blog post.
At the Google Cloud Summit in London, Google announced significant advancements in Firebase Studio, a cloud-based AI workspace designed to facilitate the development of full-stack AI applications. These updates include three versatile Agent modes in Firebase Studio for interacting with the AI model Gemini, which offers autonomous capabilities for coding tasks. Users can seamlessly toggle between modes such as Ask, Agent, and Agent (Auto-run) to brainstorm, review proposed changes, or allow Gemini to autonomously generate and modify code, all while ensuring user oversight and security. The integration of the Model Context Protocol (MCP) allows for personalized workflows, and the Gemini CLI provides a powerful, AI-driven tool for tasks beyond code, directly integrated within Firebase Studio. These innovations aim to streamline development processes and have already been applied across various industries, from hydrogen economy platforms to AI-powered fashion styling tools.
Jul 10, 2025
764 words in the original blog post.
GenAI Processors, an open-source Python library by Google DeepMind, aims to streamline the development of AI applications using Large Language Models (LLMs), particularly those requiring multimodal input and real-time responsiveness. The library introduces a consistent Processor interface that manages input handling, pre-processing, model calls, and output processing as asynchronous streams, allowing for seamless data flow and concurrency optimization. This modular approach enhances the creation of responsive applications by breaking down workflows into self-contained units and leveraging Python's asyncio for efficient task handling. GenAI Processors integrate with the Gemini API, facilitating interaction with real-time and turn-based models, and support extensive customization through extensible processors and unified multimodal handling. Currently in its early stages, the library encourages community contributions to expand its functionality and is available for exploration and development on GitHub.
Jul 10, 2025
1,169 words in the original blog post.
In the evolving field of large language models, the T5Gemma introduces a novel approach by adapting pretrained decoder-only models into encoder-decoder architectures, leveraging the Gemma 2 framework. This method, known as model adaptation, uses the weights of existing decoder-only models to initialize encoder-decoder models, subsequently refining them through UL2 or PrefixLM-based pre-training. The T5Gemma models, which include various sizes such as Small, Base, Large, XL, and unbalanced configurations like a 9B encoder with a 2B decoder, excel in inference efficiency and quality across benchmarks like SuperGLUE and GSM8K. The approach not only maintains a high quality-inference efficiency ratio but also demonstrates significant gains in tasks requiring complex reasoning, outperforming previous models in both foundational and fine-tuned capabilities. The release of T5Gemma checkpoints aims to spur further research and development within the community, offering pretrained and instruction-tuned variants to explore model architecture, efficiency, and performance.
Jul 09, 2025
922 words in the original blog post.
Gemini models have introduced a Batch Mode in their API, designed for high-throughput, non-latency-critical workloads, offering a more cost-effective solution with a 50% discount compared to synchronous APIs. This new asynchronous endpoint allows users to submit large jobs, offload scheduling and processing, and retrieve results within 24 hours, providing significant benefits such as cost savings, higher throughput, and simplified API calls without complex client-side management. Developers are leveraging Batch Mode for tasks such as bulk content generation and processing, and model evaluations, allowing for massive scalability and accelerated client deliverables. The API is designed to be intuitive, requiring users to package requests into a single file for submission and retrieval of results after job completion. Batch Mode is now available to all users, with ongoing efforts to expand its capabilities for enhanced batch processing.
Jul 07, 2025
444 words in the original blog post.