Home / Companies / Google Cloud / Blog / May 2025

May 2025 Summaries

23 posts from Google Cloud

Filter
Month: Year:
Post Summaries Back to Blog
In the modern marketing landscape, data is essential not only for measuring success but for guiding strategy, with developers playing a crucial role in transforming raw data into actionable insights. The text discusses three developer-friendly MarTech solutions that enhance the capabilities of marketing data. sGTM Pantheon offers a suite of server-side Google Tag Manager tools to manage marketing data with improved privacy, control, and performance, while GA4 Dataform transforms raw Google Analytics 4 data into accessible insights through organized, modular tables via BigQuery. Additionally, FeedX provides an open-source framework for A/B testing in Google Ads shopping campaigns, allowing advertisers to make informed decisions by comparing performance metrics through controlled experiments. These tools collectively empower developers to enhance marketing efforts by providing transparency, advanced analytics, and personalized data-driven solutions.
May 29, 2025 1,068 words in the original blog post.
The Gemini-backed Magic Mirror project transforms an ordinary mirror into an interactive gateway utilizing the Gemini API and JavaScript GenAI SDK to offer a new chat interface with real-time voice interactions, storytelling, and instant information retrieval. The mirror engages users in fluid conversations via the Live API, which allows dynamic dialogues and story weaving by utilizing the advanced generation capabilities of the Gemini model. It also incorporates Google Search for up-to-date information and uses Function Calling for image generation based on user descriptions, enhancing the storytelling and interaction experience. The project exemplifies the integration of sophisticated AI into daily life, enabling applications ranging from personalized assistants to educational and entertainment platforms, with the code and technical details available on GitHub and Hackster.io for further exploration and innovation.
May 28, 2025 637 words in the original blog post.
Google Pay is now supported within Android WebView starting with version 137, utilizing the Payment Request API to enable Android payment apps when a website is embedded inside a WebView. To implement this feature, developers must update their build dependency to androidx.webkit:webkit:1.14.0, add specific queries to their AndroidManifest.xml, and enable the Payment Request API in their WebView settings. The integration requires allowing JavaScript, enabling payment requests through WebSettings, and ensuring production access via the Pay & Wallet console. This advancement allows seamless Google Pay transactions within Android apps, enhancing the user experience by facilitating native payment processes. Developers can seek further assistance through the Google Pay & Wallet Console, engage with the developer community on Discord, or follow updates on X for ongoing support and inquiries.
May 28, 2025 394 words in the original blog post.
The Gemini API by Google provides developers with a versatile platform for creating innovative applications using advanced generative AI models, offering capabilities for text, image, and video prompts. Google AI Studio facilitates rapid prototyping and experimentation, with models like Gemini 2.5 Flash Preview showcasing improvements in reasoning, code generation, and efficiency, alongside text-to-speech functionalities supporting multiple speakers and languages. Additionally, the API introduces Lyria RealTime for live music generation, Deep Think for complex reasoning, and Gemma 3n for optimized use on everyday devices. New features include thought summaries, thinking budgets, URL context tools, computer use tools, and improvements in structured outputs and video understanding, enhancing the API's utility for AI development. Experimental capabilities like asynchronous function calling and a new Batch API are being tested, promising cost-effective and efficient data processing solutions. Overall, the Gemini API and Google AI Studio empower developers to craft diverse AI-driven applications with enhanced interactivity and performance.
May 23, 2025 1,252 words in the original blog post.
The rapid evolution of the smart home ecosystem is being driven by Gemini and the Home APIs, which aim to create seamless, intuitive experiences beyond mere device connectivity. By integrating Google's AI capabilities, the Home APIs empower developers to build innovative smart home devices and applications, expanding access to over 750 million devices, including Google's hubs and Matter infrastructure. At Google I/O 2024, the introduction of Home APIs enabled developers to enhance user experience through automation, intelligence, and connectivity, with notable partners such as ADT, LG, and Eve showcasing their innovations. Gemini-powered features offer sophisticated automation capabilities, such as suggested automations and natural language commands, enhancing user interaction with smart home devices. Developers are invited to participate in an early access program and themed Developer Challenges to explore and innovate with these new tools, while user feedback aims to shape future Google products.
May 22, 2025 1,031 words in the original blog post.
Google Wallet has expanded its availability to over 140 countries and introduced several new features aimed at enhancing user experience and engagement. Among the key updates are the integration of digital IDs, which are now accessible in multiple U.S. states and for U.K. passport holders, facilitating easier identity verification and customer experiences. The introduction of the Digital Credentials API allows apps and websites to request verifiable proof of age or identity. Google Wallet has also launched a feature enabling children to use the app with parental oversight, enhancing family usability. The app now includes advanced notification capabilities, such as field update and Nearby Passes notifications, which offer timely and relevant information when users approach specific locations. Additionally, new features like Value Added Opportunities and Pass Upgrade experiences aim to foster dynamic user interactions and seamless journeys, particularly in travel, with real-time updates and linked passes for frequent flyers. Security is bolstered with the introduction of Secure Private Images for personalized passes. These enhancements reflect Google Wallet's commitment to providing a robust platform for developers to create more engaging and secure user experiences.
May 22, 2025 1,551 words in the original blog post.
Google AI Studio has introduced a suite of new features and models designed to enhance the development of AI-powered applications, leveraging the Gemini 2.5 API. This includes native code generation with Gemini 2.5 Pro, which simplifies app creation through text, image, or video prompts, and enables developers to build and deploy web apps rapidly via the new Build tab. The platform also integrates with Google Cloud Run for seamless app deployment and utilizes a placeholder API key to manage API usage without affecting user quotas. Users can explore advanced AI models like Imagen, Veo, and Lyria RealTime for generative media and interactive music, while text-to-speech (TTS) capabilities and native audio dialog are enhanced with support for multiple voices and speaker differentiation. Additionally, the Model Context Protocol (MCP) now facilitates easier integration with open-source tools, and a new URL Context tool aids in content retrieval for tasks like fact-checking and summarization. These developments position Google AI Studio as a hub for developers to experiment with cutting-edge AI technologies.
May 21, 2025 812 words in the original blog post.
At Google I/O 2025, updates to the Google Pay API were announced, aimed at improving checkout experiences by enhancing conversion rates, boosting security, and simplifying integration processes for developers. Among the key updates are the ability to use Google Pay within Android WebViews starting with Chrome v137, improved user interfaces with richer card art, dark mode support, and greater customization options for the createButton API. The API now supports Merchant-Initiated Transactions, providing device-independent tokens and lifecycle notifications. Developers can benefit from enhanced testing environments, detailed debugging features, and new resources like codelabs and templates. Security measures have been bolstered with advanced fraud detection models and built-in identity verification, with future updates to include more detailed risk information. A new Google Pay API Status Dashboard offers real-time monitoring of API uptime and health, ensuring developers remain informed about service availability.
May 21, 2025 865 words in the original blog post.
AI agents, which perceive, decide, and act to reach specific goals, are gaining traction, with Google's Gemini models offering a robust foundation for their development. These models excel in advanced reasoning, multimodality, and function calling, crucial for creating sophisticated agents. Developers can harness open-source frameworks like LangGraph, CrewAI, LlamaIndex, and Composio to build versatile applications, each offering unique strengths. LangGraph is ideal for stateful, multi-actor workflows; CrewAI optimizes multi-agent collaboration; LlamaIndex focuses on knowledge agents through data integration; and Composio simplifies API interactions, leveraging Gemini's capabilities for real-world tasks. Selecting the appropriate framework and utilizing best practices such as iterative development and effective prompt engineering are key steps in building effective AI agents with Gemini models.
May 20, 2025 820 words in the original blog post.
At Google I/O, numerous innovations were showcased, highlighting the integration of advanced AI models, particularly from Google DeepMind, across various Google platforms. Key announcements included the introduction of Google AI Studio with the Gemini API for rapid model evaluation and building, Gemini's advanced reasoning capabilities for developing agentic experiences, and new tools like Jules, an asynchronous coding agent for GitHub repositories. The event also emphasized the enhancement of Android applications with generative AI, the introduction of new web development tools, and advancements in Firebase for building AI-powered apps. Additionally, Google introduced open models like Gemma for custom AI model training and tuning, and MedGemma for multimodal medical applications. The conference concluded with a call for developers to engage with the global community and explore further learning opportunities through various sessions and resources.
May 20, 2025 1,627 words in the original blog post.
Stitch, an experimental tool from Google Labs, aims to streamline the transition from design to development by transforming simple prompts and image inputs into complex UI designs and frontend code within minutes. This innovative tool leverages Gemini 2.5 Pro's multimodal capabilities to enhance the workflow between designers and developers, allowing for seamless collaboration. Stitch offers functionalities such as generating UI from natural language descriptions, images, or wireframes, facilitating rapid iteration and design exploration with multiple interface variants, and providing seamless integration into design systems with a paste to Figma feature. It also exports clean, functional front-end code, making the design ready for immediate development. This tool is designed to democratize app creation, making it accessible and efficient for users to build and refine applications creatively and collaboratively.
May 20, 2025 426 words in the original blog post.
Gemma 3n is a pioneering open AI model designed to deliver high-performance, low-footprint AI experiences directly on mobile devices, such as phones, tablets, and laptops. Built on a new, advanced architecture developed in collaboration with companies like Qualcomm Technologies, MediaTek, and Samsung, Gemma 3n supports real-time, multimodal AI applications, enhancing personal and private user experiences. It leverages innovations such as Per-Layer Embeddings to minimize RAM usage, allowing larger models to run efficiently on mobile hardware with a memory footprint typical of smaller models. Gemma 3n's capabilities include fast response times, privacy-first local execution, expanded multimodal understanding of audio, text, and images, and improved multilingual performance, particularly in languages like Japanese and German. It offers developers a flexible framework to dynamically adjust performance and quality, empowering new on-the-go applications capable of real-time interaction with user environments. The model is currently available for early exploration through Google AI Studio and Google AI Edge, marking a significant step in making state-of-the-art AI accessible and efficient for everyday devices.
May 20, 2025 904 words in the original blog post.
Google is advancing its vision of intelligent agents as collaborative partners by releasing significant updates across its product portfolio, focusing on development tools, management interfaces, and agent communication. Key updates include the stable release of the Python Agent Development Kit (ADK) v1.0.0 for production-ready agents and the initial release of the Java ADK v0.1.0, expanding capabilities to Java developers. The introduction of the Agent Engine UI within the Google Cloud console offers a comprehensive dashboard for managing agents, enhancing control and insight into agent performance. The Agent2Agent (A2A) protocol has been updated to version 0.2, introducing stateless interactions and standardized authentication to improve secure communications between agents. The release of the A2A Python SDK facilitates easier integration of these communication capabilities, with growing industry adoption by partners such as Auth0, Box AI, Microsoft, SAP, and Zoom, showcasing the potential for multi-agent collaboration across various platforms.
May 20, 2025 968 words in the original blog post.
Google Colab, known for providing a free cloud-hosted Jupyter Notebook environment with access to Google Cloud GPUs and TPUs, is unveiling an AI-first version, announced at Google I/O. This new iteration introduces an agentic collaborator that integrates deeply into users' workflows, enabling faster problem-solving and insight acquisition. Notable features include the Gemini 2.5 Flash-powered agentic assistance, which facilitates iterative querying, intelligent error fixing, and exploratory data analysis through the Data Science Agent (DSA). Users can interact with Colab through intuitive interfaces and receive context-aware assistance for code transformations and analytical tasks. The revamped Colab aims to enhance coding efficiency, with previous Gemini integrations demonstrating over twofold improvements, and will be gradually rolled out to users, inviting them to join the Google Labs Discord community for further engagement.
May 20, 2025 652 words in the original blog post.
Google AI Edge has expanded its support for on-device small language models (SLMs), introducing over a dozen new models, including the Gemma 3 and Gemma 3n models, available through the LiteRT Hugging Face community. Gemma 3n is the first multimodal on-device model supporting text, image, video, and audio inputs, designed for enterprise use cases where larger models can be accommodated on mobile devices. The release is complemented by the new Retrieval Augmented Generation (RAG) and Function Calling libraries, which enhance the capabilities of on-device AI by allowing for application-specific data augmentation and interactive function calling. These tools enable users to leverage models efficiently even with limited connectivity, offering transformative AI features grounded in user-relevant information. The AI Edge RAG library is currently available on Android, with plans to extend to other platforms, while the function calling library facilitates the integration of language models with application functions. Additionally, the latest quantization tools offer improved int4 post-training quantization, reducing model size and latency. The ongoing development will continue to support the latest modalities and expand functionality across platforms, with updates available through the LiteRT Hugging Face Community.
May 20, 2025 1,078 words in the original blog post.
Over the past decade, the integration of powerful accelerators like GPUs and NPUs into mobile phones has significantly enhanced the performance of AI models, offering speed increases of up to 25 times compared to CPUs and reducing power consumption by five times. Despite these benefits, developers have struggled with the complexities of interfacing with hardware-specific APIs and vendor-specific SDKs. To address these challenges, the Google AI Edge team has introduced improvements to LiteRT, including a new API that simplifies on-device ML inference, cutting-edge GPU acceleration, and NPU support developed with MediaTek and Qualcomm. These updates feature MLDrift for superior GPU performance, a uniform method for developing and deploying models on various NPUs, and an advanced TensorBuffer API that reduces memory overhead. Additionally, asynchronous execution capabilities allow for more efficient parallel processing across different processors, enhancing the responsiveness and efficiency of AI applications on mobile devices. These advancements aim to provide developers with tools to maximize AI model performance on mobile platforms, with further enhancements and broader support anticipated in the coming year.
May 20, 2025 1,410 words in the original blog post.
Keras Recommenders is a newly launched library designed to enhance digital experiences by enabling developers to create advanced recommendation systems using state-of-the-art techniques. This library offers a set of APIs with building blocks tailored for ranking and retrieval tasks, which are essential for personalized interactions in various applications, such as social media feeds and video suggestions. Compatible with JAX, TensorFlow, and PyTorch, Keras Recommenders simplifies the development of performant and accurate recommender systems by providing specialized layers, losses, and metrics. The library supports standard Keras APIs for model compilation and training configuration, and future updates will include features like the keras_rs.layers.DistributedEmbedding class for extensive embedding lookups across machines. Comprehensive documentation and examples are available on the redesigned keras.io website, and the code can be accessed on GitHub, encouraging community contributions and further development of innovative recommendation systems.
May 13, 2025 557 words in the original blog post.
The general availability of the Apigee APIM Operator introduces lightweight API Management and API Gateway capabilities to Google Kubernetes Engine (GKE) environments, facilitating API management integration with Kubernetes-like YAML for cloud-native businesses. This feature aims to reduce the conceptual and operational challenges faced by service developers and platform administrators by aligning with CNCF-standardized tooling. It empowers admins to create APIM template rules with role-based access control and offers API lifecycle management through YAML-based policies, comparable to Apigee Hybrid. The release includes key features like configuring GKE clusters to use Apigee Hybrid for API management and providing factory-built starter defaults for immediate workload tailoring. By addressing the complexity of Apigee and the need to switch toolchains, the APIM Operator is designed to simplify API management and improve developer experience. Future enhancements may include support for gRPC and GraphQL and resolving current limitations on Gateway resources and policy attachments, reflecting a commitment to evolving the platform based on customer feedback.
May 12, 2025 360 words in the original blog post.
Gemini 2.5 Pro and Gemini 2.5 Flash are the latest models from the Gemini family that represent a significant advancement in video understanding, surpassing existing models like GPT 4.1 in performance on key benchmarks. These models excel in video-to-application transformations, creating learning apps from video content, generating animations with p5.js, and retrieving specific moments using audio-visual cues. Gemini 2.5 Pro provides high accuracy in temporal reasoning and moment retrieval tasks, while Gemini 2.5 Flash offers a cost-effective solution for similar use cases. Available through Google AI Studio, the Gemini API, and Vertex AI, these models support extensive video processing with features like a 'low' media resolution parameter to handle long video contexts efficiently. The models inspire new applications in education, content creation, and interactive media, demonstrating their potential for innovative use cases.
May 09, 2025 836 words in the original blog post.
Generative AI is transforming the gaming industry by enabling developers to create dynamically evolving games that offer novel player experiences. Google, particularly through its DeepMind division, has been at the forefront of this shift, developing AI models like Gemma 3 and Gemini 2.5 that integrate advanced AI features into gaming. These models, which can be run on a single GPU or TPU, support multimodal input, extended context processing, and function calling, while also understanding over 140 languages. Google introduced the open-source Gemma Unity plugin to facilitate the integration of these models into games, as demonstrated in the collaborative sample game "Gemma Journey." Additionally, Google is partnering with gaming companies like Nazara Technologies to enhance player immersion through AI. At the Games Developer Conference, Google showcased the capabilities of its AI ecosystem, which includes tools such as Vertex AI and Google Kubernetes Engine, to support scalable and personalized gaming experiences hosted on Google Cloud.
May 09, 2025 970 words in the original blog post.
In May 2024, a significant advancement in context caching aimed at reducing repetitive context costs by 75% was introduced, and now, the Gemini API is enhancing this with a new feature: implicit caching. This feature allows developers to benefit from cache savings without setting up an explicit cache, as requests sharing a common prefix with previous ones are eligible for cache hits, dynamically providing cost savings. To optimize requests for cache hits, it's recommended to keep consistent content at the start and place variable elements like user questions at the end. The minimum request size for cache eligibility has been decreased to 1024 tokens for the 2.5 Flash model and 2048 for the 2.5 Pro model. Developers can still use the explicit caching API to ensure guaranteed savings, with the usage metadata now indicating cached tokens charged at a lower price. The company expresses enthusiasm for these developments and encourages feedback on the updates.
May 08, 2025 296 words in the original blog post.
Image Generation capabilities are now available in preview with Gemini 2.0 Flash, allowing developers to integrate conversational image generation and editing through the Gemini API in Google AI Studio and Vertex AI using the model "gemini-2.0-flash-preview-image-generation." This new version offers improvements such as better visual quality, more accurate text rendering, and reduced filter block rates compared to its experimental predecessor. Exciting functionalities include recontextualizing products in new environments, collaboratively editing images in real-time, and dynamically creating new product SKUs with text rendering. Developers can start building with these native image capabilities today, with future enhancements and expanded rate limits anticipated.
May 07, 2025 316 words in the original blog post.
Gemini 2.5 Pro Preview (I/O edition) is an advanced coding model released ahead of schedule to enhance developers' capabilities in front-end and UI development as well as fundamental coding tasks. The model, ranking first on the WebDev Arena leaderboard for its ability to create aesthetically pleasing and functional web applications, is noted for its low latency and high reliability. Its improved features include the capacity for sophisticated agentic workflows and enhanced video understanding, demonstrated by its ability to create interactive learning apps from YouTube videos. Gemini 2.5 Pro supports developers in implementing new features by generating CSS code to match design styles and helps transform concepts into functional applications with a focus on aesthetics. The updated model addresses developer feedback by reducing errors in function calling and improving trigger rates, and it remains available to developers through the Gemini API in Google AI Studio and Vertex AI for enterprise users.
May 06, 2025 727 words in the original blog post.