Home / Companies / Google Cloud / Blog / Post Details
Content Deep Dive

Introducing Gemma models in Keras

Blog post from Google Cloud

Post Details
Company
Date Published
Author
Martin Görner
Word Count
865
Company Posts That Month
12
Language
English
Hacker News Points
-
Post removed?
No
Summary

Gemma is a new family of lightweight, state-of-the-art open models available in the KerasNLP collection, compatible with JAX, PyTorch, and TensorFlow, and designed specifically for large language models. These models, offered in 2B and 7B parameter sizes, outperform similar models and some larger ones on benchmarks like MMLU, GSM8K, and HumanEval. Keras 3 introduces new features like a LoRA API for parameter-efficient fine-tuning and large-scale model-parallel training, allowing Gemma models to be fine-tuned with fewer parameters. The models can be instantiated easily with a single line of code and benefit from Keras's new distribution API for distributed training, particularly on the JAX backend due to its scalability. Users can explore various training setups, such as using TPUv3 or multiple GPUs, and are encouraged to share their fine-tuned models on platforms like Kaggle.

Trends Found in this Post
Trend Post Mentions Total Month Mentions Posts Companies MoM
AI Model Fine-tuning 6 474 91 59 +12%
TPUs 4 7 4 3 +133%
Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.