Sapling Logo

Gemma vs. Guanaco

LLM Comparison


Gemma

Gemma

Overview

Gemma is Google's family of efficient open models for local, edge, and self-hosted deployment. Gemma 4 is the current generation, spanning mobile-scale through workstation-scale models.


Google introduced Gemma in February 2024 as a family of lightweight open models built from the same research used for Gemini. Gemma 2 improved efficiency and model quality, while Gemma 3 added multimodal input and longer context. Gemma 4, released in April 2026 under Apache 2.0, adds advanced reasoning, agentic tool use, code generation, image and video understanding, and models ranging from edge-focused E2B and E4B variants to 12B, 26B mixture-of-experts, and 31B dense checkpoints.


Initial release: 2024-02-21

Current generation: Gemma 4

Guanaco

Guanaco

Overview

Guanaco is an LLM based off the QLoRA 4-bit finetuning method developed by Tim Dettmers et. al. in the UW NLP group. Guanaco achieves 99% ChatGPT performance on the Vicuna benchmark.


Guanaco is an LLM that uses a finetuning method called LoRA that was developed by Tim Dettmers et. al. in the UW NLP group. With QLoRA, it becomes possible to finetune up to a 65B parameter model on a 48GB GPU without loss of performance relative to a 16-bit model. The Guanaco model family outperforms all previously released models on the Vicuna benchmark. However, given the models are based off of the LLaMA model family, commercial use is not permitted.


Initial release: 2023-05-23

Gemma

Guanaco

Products & Features
Instruct Models
Coding Capability
Customization
Finetuning
Open Source
License Apache 2.0 Noncommercial
Model Sizes E2B, E4B, 12B, 26B (3.8B active), 31B 7B, 13B, 33B, 65B