June 2025 Summaries
12 posts from Google Cloud
Filter
Month:
Year:
Post Summaries
Back to Blog
Data Commons, an open-source knowledge graph developed by Google, aims to make publicly available statistical data more accessible and useful by unifying diverse data sources for developers, researchers, and data analysts. With the new Python client library based on the V2 REST API, developers can leverage this extensive data ecosystem more effectively, benefiting from enhancements such as Pandas dataframe APIs, integration with Pydantic libraries, and support for multiple response formats. The library's development was significantly influenced by a partnership with The ONE Campaign, exemplifying the platform's goal of fostering community contributions and innovative uses. Organizations like the United Nations can host custom Data Commons instances to integrate proprietary datasets while retaining control over their data. The new library simplifies querying the knowledge graph and supports various domains like demographics, economy, and health. Google encourages upgrading from the V1 API, which is set for deprecation, to ensure access to the latest features and support. The project highlights the power of open-source collaboration, with resources available for getting started, including documentation and tutorials.
Jun 26, 2025
660 words in the original blog post.
Gemma 3n, an advanced mobile-first AI model, marks a significant leap in on-device AI capabilities, supporting multimodal inputs and optimized for edge devices. This release builds on the success of its predecessors, featuring innovative architectures like the MatFormer for elastic inference and Per-Layer Embeddings for memory efficiency. It supports diverse inputs such as image, audio, video, and text, with enhanced features including Automatic Speech Recognition and Translation, and a state-of-the-art vision encoder, MobileNet-V5, for real-time video analysis. The model is designed to be versatile and efficient, with two sizes, E2B and E4B, offering high performance with a reduced memory footprint. Gemma 3n is supported by numerous popular AI tools and platforms, encouraging community-driven innovations through the Gemma 3n Impact Challenge, which offers substantial prizes for impactful applications. The initiative emphasizes accessibility and collaboration, aiming to inspire developers to harness the model's potential for creating transformative applications.
Jun 26, 2025
1,703 words in the original blog post.
The exploration of generative user interfaces introduces an innovative approach to human-computer interaction, where interfaces are dynamically created in real-time by a large language model like Gemini 2.5 Flash-Lite, which is crucial for maintaining low latency and responsiveness. This research prototype simulates an operating system that adapts its screens based on user actions, using a structured prompt divided into a "UI constitution" and "UI interaction" to ensure consistency while providing novel experiences. Interaction tracing enriches contextual awareness, enhancing the relevance of generated screens, and streaming techniques enable progressive rendering for a seamless user experience. To address the challenge of maintaining state, an in-memory cache is used to store session-specific UI graphs, allowing for a balance between generative flexibility and stateful consistency. Potential applications of this technology include contextual shortcuts and hybrid generative modes in existing applications, highlighting the transformative potential of generative interfaces as models evolve and improve.
Jun 25, 2025
913 words in the original blog post.
The latest generation of Gemini models, specifically the 2.5 Pro and Flash, are advancing the field of robotics with enhanced capabilities in coding, reasoning, and multimodal processing, combined with spatial understanding. These models enable developers to create sophisticated robotics applications by utilizing features such as semantic scene understanding, multimodal reasoning, and spatial reasoning integrated with code generation for robot control. Gemini 2.5 can perform complex tasks like identifying objects in camera feeds, understanding and responding to voice commands, and generating robot control codes for tasks such as moving objects. The Live API further facilitates real-time interactive applications, allowing voice control over robots through function calls. These models have shown robust performance on benchmarks, ensuring safety and preventing violations of ethical and safety policies. The Gemini Robotics-ER model, released earlier in March, has already inspired various applications by companies like Agile Robots and Boston Dynamics, showcasing the potential of these models in robotics development.
Jun 24, 2025
2,217 words in the original blog post.
KerasHub is a Python library that enhances the flexibility of defining and utilizing machine learning models by allowing the integration of popular model architectures and their weights across different machine learning frameworks such as JAX, PyTorch, and TensorFlow. It supports interoperability with repositories like the Hugging Face Hub, where models are often saved in the SafeTensors format, enabling users to load these checkpoints into KerasHub models irrespective of the original framework used to create them. This capability allows users to mix and match model architectures with different sets of weights, facilitating experimentation and innovation without being confined to a single ecosystem. KerasHub simplifies the process by offering built-in converters for Hugging Face transformer models, ensuring seamless loading of a wide variety of pretrained models into KerasHub with minimal code. This flexibility empowers users to leverage a vast collection of community fine-tuned models while retaining the freedom to choose their preferred backend framework, thus bridging the gap between different frameworks and checkpoint repositories.
Jun 24, 2025
1,625 words in the original blog post.
Imagen 4, the latest text-to-image model from Google, is now available for paid preview via the Gemini API and limited free testing in Google AI Studio. This new model family, which includes Imagen 4 and Imagen 4 Ultra, marks a significant advancement in text rendering and image generation quality. Imagen 4 is designed for a wide range of tasks and is priced at $0.04 per output image, while Imagen 4 Ultra offers more precise alignment with text prompts at $0.06 per output image. Upcoming billing tiers and higher rate limits are planned. Examples of Imagen 4 Ultra's capabilities include diverse styles and content, such as cosmic comic panels, vintage travel postcards, and fashion editorials. All images generated include a non-visible digital SynthID watermark to ensure trust and transparency. Users can explore the models further with official documentation and cookbooks as they await broader availability.
Jun 24, 2025
455 words in the original blog post.
At Google I/O 2025, Google announced the public release of an AI-first version of Colab, aimed at becoming a comprehensive coding partner within users' notebooks. Initially available to a small group, the new features received positive feedback for aiding users in accelerating projects, learning new skills, and gaining insights from data. The AI-first Colab supports end-to-end machine learning projects by autonomously handling tasks from data preparation to model evaluation, enhances debugging by acting as a pair programmer that suggests fixes, and simplifies data visualization by generating high-quality charts. Key features include iterative querying for conversational code requests, a Next-Generation Data Science Agent for autonomous analytical workflows, and effortless code transformation through natural language descriptions. Google encourages users to explore these capabilities by accessing any Colab notebook and engaging with the AI through the Gemini spark icon, while also inviting feedback and community interaction via their Google Labs Discord channel.
Jun 24, 2025
514 words in the original blog post.
The Unlock Global Communication with Gemma competition on Kaggle showcased the innovative efforts of developers in adapting large language models (LLMs) to various cultural and linguistic contexts, addressing biases towards high-resource languages. Participants creatively tackled translation of languages, lyrics, old texts, and more, using custom datasets and efficient post-training methods for instruction following and domain-specific tasks. Notable projects included adapting Gemma for Swahili, Arabic, Traditional Chinese, Italian, Ancient Chinese, and other languages, demonstrating the potential of LLMs for cultural preservation and educational tools. The competition highlighted the community's dedication to expanding the capabilities of AI for underrepresented languages, setting the stage for the upcoming Gemma 3, which will support over 140 languages, further bridging communication gaps worldwide.
Jun 23, 2025
1,138 words in the original blog post.
Open Source Summit North America witnessed the Linux Foundation's announcement of the Agent2Agent (A2A) project, a collaborative effort among major tech companies like Amazon Web Services, Cisco, Google, Microsoft, Salesforce, SAP, and ServiceNow, aimed at fostering an open and interoperable AI agents ecosystem. This initiative, hosted by the Linux Foundation, seeks to break down current silos in AI through the A2A protocol, an open standard designed to enable communication and collaboration between different AI agents. With over 100 companies supporting this protocol, A2A offers a common language for AI agents to discover capabilities, exchange information securely, and coordinate complex tasks, paving the way for more powerful and innovative AI applications. The Linux Foundation's neutral governance ensures that the A2A project remains vendor-agnostic and community-driven, accelerating the protocol's adoption by providing a robust framework for collaboration and intellectual property management. Founding members express their commitment to integrating A2A into their platforms to enable seamless interoperability and enhanced digital workforce solutions, with the project inviting global participation to contribute to the future of AI.
Jun 23, 2025
887 words in the original blog post.
Gemini Code Assist is now generally available in Apigee API Management, offering AI-assisted API development capabilities designed to streamline the creation of consistent, secure, and well-designed APIs. This tool leverages Google's Gemini models and Apigee's Enterprise Context to align generated APIs with organizational standards, ensuring consistency and security. Key features include a chat interface for API creation, AI-generated specification summaries, iterative spec design, duplicate API detection, and enterprise-grade security compliance. The tool facilitates a seamless development workflow by allowing developers to generate, iterate, test, and implement APIs using natural language prompts, thus reducing duplication and enhancing governance. Existing customers can access these capabilities through VS Code, with future expansions planned for additional IDE support and enhanced functionalities.
Jun 18, 2025
638 words in the original blog post.
Gemini 2.5 introduces a family of enhanced thinking models, including the Gemini 2.5 Pro, Flash, and Flash-Lite, each designed to improve performance and accuracy through dynamic control over their thinking budgets. The Gemini 2.5 Flash-Lite model, now available in preview, offers the lowest latency and cost, making it suitable for high-throughput tasks like classification and summarization, while supporting native tools such as Google Search grounding and code execution. Pricing updates have been made for the Gemini 2.5 Flash, removing the distinction between thinking and non-thinking pricing and adjusting costs to reflect its enhanced value. The Gemini 2.5 Pro, known for its high intelligence and capability in demanding tasks, is now stable and continues to see significant growth and demand. As these models become more available, developers are offered more flexibility and cost-effective options for deploying these advanced AI solutions.
Jun 17, 2025
677 words in the original blog post.
PCI DSS v4 compliance for checkout pages involves managing payment scripts with authorization, integrity assurance, and inventory maintenance. While techniques like Subresource Integrity (SRI) aren't feasible for Google Pay's pay.js due to its build process, using a sandboxed iframe satisfies compliance by isolating scripts from the parent DOM. This approach, which involves specific sandbox attribute values such as allowing scripts, popups, same-origin access, and forms, has been successfully implemented by Shopify, enabling them to pass the PCI DSS v4 audit. By integrating Google Pay within a sandboxed iframe, businesses can maintain secure and compliant checkout processes, and further support can be sought through the Google Pay & Wallet Console or developer community channels.
Jun 10, 2025
502 words in the original blog post.