Gemma 4 31B
Gemma 4 31B is Google DeepMind’s largest Gemma 4 open-weight dense multimodal model, featuring around 31 billion parameters and strong performance on text and image understanding tasks.
What is Gemma 4 31B?
Gemma 4 31B is a 31-billion-parameter dense multimodal large language model from Google DeepMind that processes text and images with text outputs. It is primarily used for advanced assistant-style chat, coding help, and analytical reasoning tasks that benefit from long-context understanding. It is also applied to multimodal use cases such as image-grounded question answering and document understanding where both text and images must be interpreted together. It belongs to the Gemma 4 family of open models, which span multiple sizes from edge-oriented variants to this largest 31B configuration.
Providers
Different companies host the same model. OpenRouter routes your request to one of them based on the routing mode you pick — Balanced (price + speed), Nitro (fastest), or Exacto (highest tool-calling accuracy).
| Provider | Input | Output | Cache read /M | Latency | Throughput | Uptime |
|---|---|---|---|---|---|---|
| ~$0.35 per 1M tokens | ~$0.70 per 1M tokens | — | ~220ms | ~80 tps | 100.00% | |
| Vertex AI (Google Cloud) | ~$0.38 per 1M tokens | ~$0.76 per 1M tokens | — | ~260ms | ~60 tps | ~99.9% |
| AWS Bedrock (3rd‑party Gemma‑equivalent) | ~$0.40 per 1M tokens | ~$0.80 per 1M tokens | — | ~250ms | ~70 tps | ~99.9% |
| Anthropic (Claude Sonnet‑class alternative) | ~$0.50 per 1M tokens | ~$1.00 per 1M tokens | — | ~230ms | ~75 tps | ~99.9% |
Try this model
Test Gemma 4 31B right here — free to start.
Suggestions for your first prompt
Code snippet
Call the model through the OpenAI-compatible API.
from openai import OpenAI
client = OpenAI(
api_key="YOUR_API_KEY",
base_url="https://inference.example.com/v1"
)
response = client.chat.completions.create(
model="google/gemma-4-31b",
messages=[
{
"role": "user",
"content": "Describe this image in one sentence."
}
],
)
print(response.to_json())
{
"model": "google/gemma-4-31b",
"messages": [
{
"role": "user",
"content": "Describe this image in one sentence."
}
]
}
Uptime
Last 30 days
30/30 days operational | 100.00% uptime
5 Core Capabilities
-
Advanced Reasoning
Performs complex, step-by-step reasoning for difficult tasks, benefiting from an explicit thinking mode in instruction-tuned variants.
-
Multimodal Understanding
Processes text and images together, supporting tasks like document parsing, UI comprehension, charts, and general visual understanding.
-
Conversational Chat
Acts as a strong conversational assistant, following instructions, maintaining context, and supporting agentic workflows and tool use.
-
Code Generation
Generates, completes, and debugs source code in multiple languages, suitable for software development and technical scripting tasks.
-
Multilingual Text
Handles multilingual input and output across many languages, enabling translation-style tasks and cross-lingual reasoning over long context.
6 Most Valuable Use Cases
- Customer Support Chatbots
- Invoice Data Extraction
- Legal Document Review
- Compliance Case Monitoring
- E-commerce Product Assistants
- Code Generation Assistance
Why Build on LLM.API?
One unified API. Every major model. Built-in reliability, cost control, and observability.
-
Intelligent Model Routing
Automatically route each request to the optimal model across providers based on latency, cost, and capability — no client changes required.
One endpoint, every model -
Cost-Aware Orchestration
Optimize spend with dynamic model selection, rate limiting, and usage controls that keep your AI bill predictable while preserving performance.
Lower cost, same quality -
Resilient Fallback Logic
Define cross-provider failover rules so requests automatically retry on backup models when a provider is down, slow, or throttling.
No single point of failure -
End-to-End Observability
Get unified logs, metrics, traces, and payload sampling across all providers to debug failures, tune prompts, and monitor performance in one place.
See every token, everywhere -
Task-Level Abstractions
Call high-level tasks like chat, RAG, tools, or agents without wiring each provider’s primitives yourself, so you ship features instead of glue code.
APIs speak in tasks -
High-Throughput Batch Jobs
Run large-scale inference workloads with parallel execution, retries, and progress tracking built in, without manually managing queues or worker pools.
Scale from 10 to 10M
When to Use — When NOT to Use
Use it if...
- You need a strong open-weight LLM that can be self-hosted on your infrastructure.
- You need high-quality English and multilingual text generation for chatbots or virtual assistants.
- Your use case involves fine-tuning or LoRA adapters on a powerful 30B-class backbone.
- Your use case involves moderate-length coding help, code explanations, and boilerplate generation.
- You need a balance of reasoning quality and cost compared with much larger proprietary models.
- Your use case involves RAG over medium documents where ultra-long context is unnecessary.
- You need an open model compatible with common inference stacks like vLLM or Ollama.
Avoid if...
- You need cutting-edge reasoning and tool use rivaling the very best flagship proprietary models.
- Your workload requires extremely long-context processing, such as full-book analysis or codebases.
- You need highly optimized edge or mobile deployment where a 31B model is impractical.
- You need top-tier, production-grade code synthesis for complex multi-file or large refactor tasks.
- Your workload requires guaranteed low-latency responses on modest GPUs or CPU-only environments.
- You need native, fully managed hosting with tight integration into non-Google cloud ecosystems.
- Your workload requires robust, battle-tested safety layers and policy enforcement out-of-the-box.
Frequently Asked Questions
-
What is Gemma 4 31B?
Gemma 4 31B is a 31-billion-parameter Google language model focused on strong reasoning, coding, and instruction-following capabilities via the LLM.API gateway.
-
What is the context window of Gemma 4 31B?
Gemma 4 31B supports a 32K token context window, allowing relatively long conversations and documents before older tokens are pushed out.
-
What is Gemma 4 31B best suited for?
Gemma 4 31B is best for complex reasoning, multi-step agents, advanced coding assistance, and high-quality English writing where accuracy matters.
-
Does Gemma 4 31B support images or other modalities?
Gemma 4 31B is a text-only model on LLM.API, supporting text inputs and outputs but not images, audio, or video.
-
How fast is Gemma 4 31B when called through LLM.API?
On LLM.API, Gemma 4 31B typically returns first tokens within a few hundred milliseconds and then streams tokens at an interactive rate.
-
How is Gemma 4 31B priced on LLM.API?
Gemma 4 31B pricing on LLM.API is usage-based per input and output token; check the LLM.API pricing page for up-to-date rates.
-
How do I call Gemma 4 31B via LLM.API?
You select the Gemma 4 31B model name in your LLM.API request and authenticate with your LLM.API key, without needing direct Google Cloud setup.
-
How does Gemma 4 31B compare to smaller Gemma variants?
Compared to smaller Gemma models, Gemma 4 31B generally offers better reasoning quality and coding ability at the cost of higher latency and price.
-
What are the main limitations of Gemma 4 31B?
Gemma 4 31B can hallucinate facts, lacks real-time web access, may underperform on niche domains, and is restricted to its context window.
-
Can I use Gemma 4 31B for structured outputs like JSON?
Yes, Gemma 4 31B can reliably follow JSON or schema-like formats when prompted clearly and validated by your application logic.
COMPARE
Competitive Models
-
Nano Banana (Gemini 2.5 Flash Image)
Nano Banana (Gemini 2.5 Flash Image) is Google’s high-speed visual generation and editing model designed for low-latency, high‑volume image workflows with strong character and style consistency.
-
Gemini 3.5 Flash
Gemini 3.5 Flash is Google’s natively multimodal reasoning model optimized for very low latency and cost while maintaining frontier‑level performance, particularly for coding and agentic workflows.
-
Nano Banana 2 (Gemini 3.1 Flash Image Preview)
Nano Banana 2 (Gemini 3.1 Flash Image Preview) is Google DeepMind’s image generation and editing model built on the Gemini 3.1 Flash architecture, optimized for fast, cost‑efficient, high‑quality visuals. It balances strong multimodal understanding with 4K-capable output and low latency for both text-to-image and image-edit tasks.
Get one key to every model
Swap your API key. Keep your code.