Nano Banana Pro (Gemini 3 Pro Image Preview)
Nano Banana Pro (Gemini 3 Pro Image Preview) is Google’s preview-stage image generation and editing model built on the Gemini 3 Pro family, optimized for complex, multi-turn visual creation tasks.
What is Nano Banana Pro (Gemini 3 Pro Image Preview)?
Nano Banana Pro (Gemini 3 Pro Image Preview) is a proprietary Google Gemini 3 image model that generates and edits images from text prompts and reference images. It is mainly used for high-quality, multi-step image creation and editing workflows, such as photorealistic rendering, design mockups, and creative compositing. It also supports multimodal use cases where text and images are combined, leveraging a context window of around 65k tokens for detailed, instruction-heavy prompts. The model belongs to the Gemini 3 family and is a preview/legacy variant of the Gemini 3 Pro Image line, with newer Nano Banana Pro image models recommended for new integrations.
Providers
Different companies host the same model. OpenRouter routes your request to one of them based on the routing mode you pick — Balanced (price + speed), Nitro (fastest), or Exacto (highest tool-calling accuracy).
| Provider | Input | Output | Cache read /M | Latency | Throughput | Uptime |
|---|---|---|---|---|---|---|
| $2.00 | $12.00 | $0.20 | ~300ms | ~60 img/min | 100.00% | |
| $2.00 | $12.00 | $0.20 | ~300ms | ~60 img/min | 100.00% |
Try this model
Test Nano Banana Pro (Gemini 3 Pro Image Preview) right here — free to start.
Suggestions for your first prompt
Code snippet
Call the model through the OpenAI-compatible API.
from openai import OpenAI
client = OpenAI(
api_key="YOUR_API_KEY",
base_url="https://inference.example.com/v1"
)
response = client.chat.completions.create(
model="google/nano-banana-pro-gemini-3-pro-image-preview",
messages=[
{
"role": "user",
"content": "Describe this image in one sentence."
}
],
)
print(response.to_json())
{
"model": "google/nano-banana-pro-gemini-3-pro-image-preview",
"messages": [
{
"role": "user",
"content": "Describe this image in one sentence."
}
]
}
Uptime
Last 30 days
30/30 days operational | 100.00% uptime
5 Core Capabilities
-
Natural Conversation
Engages in multi-turn dialogue, answering questions and following instructions with context awareness across a wide range of topics.
-
Image Interpretation
Analyzes user-provided images to recognize objects, scenes, and visual relationships, supporting grounded reasoning about visual content.
-
Text Translation
Translates text between multiple languages, preserving meaning and tone for everyday communication and informational content.
-
Visual Text Extraction
Extracts readable text from images, enabling recognition of signs, documents, labels, and other embedded text in pictures.
-
Tool Integration
Coordinates with external tools or systems, using model outputs to support monitoring, analysis, or automated workflows.
6 Most Valuable Use Cases
- Mobile Vision Inference
- On-device Image Captioning
- AR Object Detection
- Robotics Scene Understanding
- Smart Camera Automation
- Privacy-preserving Analytics
Why Build on LLM.API?
One unified API. Every major model. Built-in reliability, cost control, and observability.
-
Unified AI Routing
Dynamically route each request to the optimal model across providers based on latency, cost, or performance policies—without changing your application code.
One endpoint, any model -
Cost-Aware Orchestration
Automatically balance premium and budget models using configurable cost guards, so you control spend while keeping response quality and performance predictable at scale.
Optimize spend by default -
Resilient Fallbacks
Define automatic cross-provider fallbacks so your workloads keep running through rate limits, outages, or model deprecations—with no manual error-handling sprawl.
Stay online, by default. -
Deep Observability
Get unified logs, metrics, traces, and request payloads across all models and vendors, with per-model performance and cost breakdowns for real-time debugging and tuning.
See every token, everywhere. -
Task-Level Abstractions
Describe tasks like chat, generation, tools, or RAG once, and let LLM.API handle model-specific parameters, prompts, and formats for each provider.
Think tasks, not models. -
High-Throughput Batching
Submit large batches of prompts through one API, and let LLM.API parallelize, retry, and aggregate responses for consistent throughput across providers.
Scale tokens, not code.
When to Use — When NOT to Use
Use it if...
- You need fast, low-cost image understanding for previews, thumbnails, or basic tagging.
- You need to quickly classify or caption user-uploaded photos before further processing.
- Your use case involves simple multimodal prompts combining short text with a single image.
- Your use case involves prototyping lightweight visual features without needing top-tier reasoning.
- You need a small model to pre-filter or route images for larger backends.
- You need to extract obvious objects, colors, or layouts from everyday consumer images.
Avoid if...
- You need state-of-the-art vision-language reasoning on complex diagrams, charts, or scientific images.
- Your workload requires long-context multimodal analysis spanning many images and extensive text.
- You need highly reliable domain-specific medical or industrial image interpretation with safety guarantees.
- You need robust handling of very high-resolution images or detailed small-object detection.
- Your workload requires consistent, top-tier general reasoning or coding beyond basic visual tasks.
- You need full general-purpose chat capabilities rather than focused image preview understanding.
Frequently Asked Questions
-
What is Nano Banana Pro (Gemini 3 Pro Image Preview)?
Nano Banana Pro (Gemini 3 Pro Image Preview) is a lightweight Gemini-3–based multimodal model from Google optimized for fast image understanding and text generation.
-
What is Nano Banana Pro (Gemini 3 Pro Image Preview) best suited for?
It is best for low-latency applications like UI assistants, rapid image captioning, and lightweight reasoning over images and short texts.
-
What is the context window of Nano Banana Pro (Gemini 3 Pro Image Preview)?
Nano Banana Pro (Gemini 3 Pro Image Preview) supports a 32K token context window for combined input and output.
-
Which modalities does Nano Banana Pro (Gemini 3 Pro Image Preview) support via LLM.API?
It supports text input and output plus image input, including multi-image prompts, but does not generate images.
-
How is Nano Banana Pro (Gemini 3 Pro Image Preview) priced on LLM.API?
Pricing is per input and output token, with Nano Banana Pro positioned as a cheaper Gemini tier; check LLM.API pricing docs for current rates.
-
How fast is Nano Banana Pro (Gemini 3 Pro Image Preview) in terms of latency?
It is optimized for low latency, typically returning short responses in a few hundred milliseconds under normal load.
-
How do I call Nano Banana Pro (Gemini 3 Pro Image Preview) through the LLM.API?
Specify the provider as Google and the model name as "nano-banana-pro-gemini-3-pro-image-preview" in your LLM.API completion or chat request.
-
How does Nano Banana Pro (Gemini 3 Pro Image Preview) compare to larger Gemini models?
It trades some reasoning depth and long-context performance for significantly lower cost and faster responses than flagship Gemini 3 Pro models.
-
What limitations should I be aware of with Nano Banana Pro (Gemini 3 Pro Image Preview)?
It can hallucinate, struggles with very long multi-step reasoning or domain-expert tasks, and should not be used without human review for high-risk decisions.
-
Can I fine-tune or customize Nano Banana Pro (Gemini 3 Pro Image Preview) via LLM.API?
Direct fine-tuning is not supported; use system prompts, few-shot examples, and retrieval to adapt behavior.
COMPARE
Competitive Models
-
Gemini 3.5 Flash
Gemini 3.5 Flash is Google’s natively multimodal reasoning model optimized for very low latency and cost while maintaining frontier‑level performance, particularly for coding and agentic workflows.
-
Nano Banana (Gemini 2.5 Flash Image)
Nano Banana (Gemini 2.5 Flash Image) is Google’s high-speed visual generation and editing model designed for low-latency, high‑volume image workflows with strong character and style consistency.
-
Gemini 3.1 Flash Lite
Gemini 3.1 Flash Lite is Google’s ultra-fast, low-cost Gemini 3-series language model optimized for high-volume, latency-sensitive applications. It prioritizes speed and cost-efficiency while still supporting multimodal understanding and configurable reasoning depth.
Get one key to every model
Swap your API key. Keep your code.