Introducing Gemma models in Keras
Blog post from Google Cloud
Gemma is a new family of lightweight, state-of-the-art open models available in the KerasNLP collection, compatible with JAX, PyTorch, and TensorFlow, and designed specifically for large language models. These models, offered in 2B and 7B parameter sizes, outperform similar models and some larger ones on benchmarks like MMLU, GSM8K, and HumanEval. Keras 3 introduces new features like a LoRA API for parameter-efficient fine-tuning and large-scale model-parallel training, allowing Gemma models to be fine-tuned with fewer parameters. The models can be instantiated easily with a single line of code and benefit from Keras's new distribution API for distributed training, particularly on the JAX backend due to its scalability. Users can explore various training setups, such as using TPUv3 or multiple GPUs, and are encouraged to share their fine-tuned models on platforms like Kaggle.
| Trend | Post Mentions | Total Month Mentions | Posts | Companies | MoM |
|---|---|---|---|---|---|
| AI Model Fine-tuning | 6 | 474 | 91 | 59 | +12% |
| TPUs | 4 | 7 | 4 | 3 | +133% |
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.