Sapling Logo

Guanaco vs. Llama

LLM Comparison


Guanaco

Guanaco

Overview

Guanaco is an LLM based off the QLoRA 4-bit finetuning method developed by Tim Dettmers et. al. in the UW NLP group. Guanaco achieves 99% ChatGPT performance on the Vicuna benchmark.


Guanaco is an LLM that uses a finetuning method called LoRA that was developed by Tim Dettmers et. al. in the UW NLP group. With QLoRA, it becomes possible to finetune up to a 65B parameter model on a 48GB GPU without loss of performance relative to a 16-bit model. The Guanaco model family outperforms all previously released models on the Vicuna benchmark. However, given the models are based off of the LLaMA model family, commercial use is not permitted.


Initial release: 2023-05-23

Llama

Llama

Overview

Llama is Meta's open-weight model family. Llama 4 was its last major generation; Meta's active assistant-model development has shifted to Muse.


Meta introduced LLaMA in February 2023, helping catalyze the modern open-weight model ecosystem. Llama 2 added commercial use and chat tuning, while the Llama 3 series improved scale, multilingual support, and instruction following. Llama 4, released in April 2025, introduced native multimodality and a mixture-of-experts architecture. Meta subsequently shifted its actively promoted assistant-model line to the proprietary Muse family, but Llama models remain widely used and deployed.


Initial release: 2023-02-24

Current generation: Llama 4

Guanaco

Llama

Products & Features
Instruct Models
Coding Capability
Customization
Finetuning
Open Source
License Noncommercial Llama Community License
Model Sizes 7B, 13B, 33B, 65B 109B (17B active), 400B (17B active)