Skip to content

Evaluation

CPU inference and quantization quality evaluation

CPU-only inference speed on four architectures with gemma.cpp (8-bit switched-floating-point quantization) using a PaliGemma 2 3B (224px2) checkpoint fine-tuned on COCOcap; quality comparison between Jax/f32 inference on TPU and quantized gemma.cpp inference on CPU for five fine-tuning datasets (chosen for task coverage)

Multimodal Google DeepMind — Gemma

How it was run

Open a release to see the setup, scoring, and other details that the lab published. Only fields the lab actually disclosed are shown.

PaliGemma 2
Google DeepMind — Gemma

rich detail 28 method fields published

See how this test was run