Gemma 4 26B A4B
Googlehuggingface.co/google/gemma-4-26B-A4B-it
Gemma 4 26B A4B IT is an instruction-tuned Mixture-of-Experts (MoE) model from Google DeepMind. Despite 25.2B total parameters, only 3.8B activate per token during inference — delivering near-31B quality at...
Updated 6 h ago · first seen 11 Sept 2026
model_01M294WW1FFYS3E90RJS36V8HP
Overview
Identity
Identity block not returned by the API for this entity.
Openness
Openness not classified yet — no sourced evidence to place this model in the ontology.
Key facts
- Release date
Source:Hugging Face Hub (public pages, model cards, papers)T2observed 7 h agomedium
- Model card
Source:Hugging Face Hub (public pages, model cards, papers)T2observed 12 h agomedium
- Openrouter id
Source:OpenRouter public model & pricing listingT2observed 12 h agomedium
Architecture
- Architecture
Source:Hugging Face Hub (public pages, model cards, papers)T2observed 11 h agomedium
- Model type
Source:Hugging Face Hub (public pages, model cards, papers)T2observed 11 h agomedium
- Parameters
Source:Hugging Face Hub (public pages, model cards, papers)T2observed 12 h agomedium
- Active parameters
Source:Hugging Face Hub (public pages, model cards, papers)T2observed 11 h agomedium
- Tokenizer
Source:OpenRouter public model & pricing listingT2observed 12 h agomedium
- Weights dtype
Source:Hugging Face Hub (public pages, model cards, papers)T2observed 11 h agomedium
- File size
Source:Hugging Face Hub (public pages, model cards, papers)T2observed 11 h agomedium
- Library name
Source:Hugging Face Hub (public pages, model cards, papers)T2observed 11 h agomedium
- Pipeline tag
Source:Hugging Face Hub (public pages, model cards, papers)T2observed 12 h agomedium
- Hugging Face repo
Source:OpenRouter public model & pricing listingT2observed 12 h agomedium
Capabilities
Modalities
- Modalities
- imagetextvideo
- Input
- imagetextvideo
- Output
- text
Capabilities
Tool calling
Yes
OpenRouter public model & pricing listing · T2
Structured output
Yes
OpenRouter public model & pricing listing · T2
Reasoning
Yes
OpenRouter public model & pricing listing · T2
Vision
Yes
OpenRouter public model & pricing listing · T2
Audio
Unavailable
Fine-tuning available
Unavailable
- Context window
Source:OpenRouter public model & pricing listingT2observed 9 h agomedium
- Max output
Source:OpenRouter public model & pricing listingT2observed 12 h agomedium
- Tokenizer
Source:OpenRouter public model & pricing listingT2observed 12 h agomedium
Benchmarks8
Compare with another model →Comparable same task and conditions · Partially comparable same task, conditions differ (effort, temperature, judge) · Not comparable different variant or metric
No benchmark results recorded
Providers & Pricing2
All offers in the price terminal →USD per 1M tokens as published by each provider (USD). Rows are append-only: every change is kept in the history below.
Price history
Output price · USD / 1M tokens 1 provider
- Google Gemini API
- Google Gemini API$0.22 → $011 Sept 2026
- Google Gemini APIfirst observed $0.2211 Sept 2026
Input price · USD / 1M tokens 1 provider
- Google Gemini API
- Google Gemini API$0.042 → $011 Sept 2026
- Google Gemini APIfirst observed $0.04211 Sept 2026
Hardware fit37
Assumptions (6)
- Estimated, not measured: weights = parameters × bytes/param × 1.15 runtime overhead.
- bytes/param: 4bit = 0.5, 8bit = 1.0, fp16 = 2.0 (uniform quantization, no per-layer exceptions).
- KV cache approximated at 0.5 GB per 8 192 tokens of context, independent of architecture (GQA/MLA models need less).
- A model 'fits' when the estimate is at most the device memory minus 2 GB reserved for the OS and framework.
- Mixture-of-experts models are estimated on total parameters (all experts must be resident); active parameters are ignored.
- Device memory uses the largest configuration when several are listed (e.g. Apple silicon tiers).
Lineage
Open in Graph →- ancestor: gemma-4-26B-A4B
Papers1
- arXiv:2607.02770Active35
Timeline15
Full timeline →Gemma 4 26B A4B: context length changed from 256000 to 262144
Context window256K tokens→262.1K tokensopenrouterGemma 4 26B A4B scores 13.64% on Terminal-Bench
artificial_analysisGemma 4 26B A4B scores 38.95% on Terminal-Bench
artificial_analysisGemma 4 26B A4B scores 43.57% on τ²-bench
artificial_analysisGemma 4 26B A4B scores 69.25% on MMMU-Pro
artificial_analysisGemma 4 26B A4B scores 72.45% on IFBench
artificial_analysisGemma 4 26B A4B scores 19.32% on Humanity's Last Exam
artificial_analysisGemma 4 26B A4B scores 16.67 on Artificial Analysis Intelligence Index
artificial_analysisGemma 4 26B A4B: context length changed from 262144 to 256000
Context window262.1K tokens→256K tokensartificial_analysisGemma 4 26B A4B: release date changed from 2026-03-11 to 2026-04-03
Release date11 Mar 2026→3 Apr 2026openrouterGemma 4 26B A4B: release date changed from 2026-04-03 to 2026-03-11
Release date3 Apr 2026→11 Mar 2026huggingfaceGoogle Gemini API lists Gemma 4 26B A4B at $0 in / $0 out per 1M tokens
openrouterGoogle Gemini API lists Gemma 4 26B A4B at $0.042 in / $0.22 out per 1M tokens
openrouter
Change history42
Viewing AI Atlas as of 12 Sept 2026 — attributes exactly as the atlas knew them on that day; later corrections are not shown.
Back to today →Attributes as of 12 Sept 2026 43 claims in force
- Release date
- 11 Mar 2026
- Openness
- open-weights
- License
- Apache-2.0
- Architecture
- Gemma4ForConditionalGeneration
- Parameters
- 25.8B
- Active parameters
- 4B
- Context window
- 262.1K tokens
- Max output
- 32.8K tokens
- Modalities
- image, text, video
- Input modalities
- image, text, video
- Output modalities
- text
- Tokenizer
- Gemma
- Base model
- google/gemma-4-26B-A4B
- File size
- 51.6 GB
- Hugging Face repo
- google/gemma-4-26B-A4B-it
- Pipeline tag
- image-text-to-text
- Model card
- Downloads
- 8,983,063
- Likes
- 1,489
- Aa context window
- 256,000
- Aa openness
- open-weights
- Commercial use allowed
- Yes
- Derivatives allowed
- Yes
- Gated
- No
- Hf inference providers
- deepinfra, featherless-ai, novita, scaleway
- Last modified
- 2026-07-20T16:42:12+00:00
- Library name
- transformers
- License url
- https://ai.google.dev/gemma/docs/gemma_4_license
- Aa median output tokens per second
- 75.4
- Downloads all time
- 58,013,646
- Model type
- gemma4
- Openrouter id
- google/gemma-4-26b-a4b-it
- Openrouter listed at
- 3 Apr 2026
- Reasoning
- Yes
- Redistribution allowed
- Yes
- Structured output
- Yes
- Supported parameters
- frequency_penalty, include_reasoning, logit_bias, logprobs, max_tokens, min_p, presence_penalty, reasoning, repetition_penalty, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_k, top_logprobs, top_p
- Tags
- transformers, safetensors, gemma4, image-text-to-text
- Tool calling
- Yes
- Vision
- Yes
- Weights available
- Yes
- Weights dtype
- BF16
Release daterelease_date6
Opennessopenness1
Licenselicense1
Architecturearchitecture1
Parametersparameter_count1
Active parametersactive_parameter_count1
Context windowcontext_length3
Max outputmax_output_tokens1
Modalitiesmodalities1
Input modalitiesmodalities_input1
Output modalitiesmodalities_output1
Tokenizertokenizer1
Base modelbase_model1
File sizefile_size_gb1
Hugging Face repohf_repo1
Pipeline tagpipeline_tag1
Model cardmodel_card_url1
Downloadsmetric.downloads1
Likesmetric.likes1
Descriptiondescription1
Gatedgated1
Hf inference providershf_inference_providers1
Last modifiedlast_modified1
Library namelibrary_name1
License urllicense_url1
Downloads all timemetric.downloads_all_time1
Model typemodel_type1
Openrouter idopenrouter_id1
Reasoningreasoning1
Structured outputstructured_output1
Supported parameterssupported_parameters1
Tool callingtool_calling1
Visionvision1
Weights dtypeweights_dtype1
Claims are temporal and append-only: a new observation closes the previous claim (valid_to) instead of overwriting it. Conflicting claims from different sources are kept side by side and flagged — never averaged. Methodology →
Provenance
Attributed facts
35
Source tiers
T235
Freshest observation
6 h ago
Conflicts
None
Source documents 7
Tier 1 = official/primary, 2 = quality secondary, 3 = community, 4 = unverified. Every snapshot is archived; see all sources and the methodology.
Data quality (65/100) measures how well AI Atlas knows this entity — completeness, primary-source ratio, freshness, conflicts — never how good the model is. Methodology →