Bonus: Top up now and we'll double your first deposit — get x2 credits instantly.

Chirp 3

Chirp 3 is Google's latest-generation multilingual speech and audio model, available through Google Cloud for high-accuracy transcription and natural-sounding text-to-speech.

What is Chirp 3?

Chirp 3 is a multilingual Automatic Speech Recognition and audio generation model from Google that powers Speech-to-Text and Text-to-Speech capabilities in Google Cloud. It is used for accurate real-time and batch audio transcription across many languages, including support for speaker diarization and language-agnostic transcription. It is also used to generate high-fidelity synthetic speech, including instant custom voice models built from high-quality recordings. Chirp 3 succeeds earlier Chirp models as part of Google’s Chirp family of speech and audio foundation models.


Providers

Different companies host the same model. OpenRouter routes your request to one of them based on the routing mode you pick — Balanced (price + speed), Nitro (fastest), or Exacto (highest tool-calling accuracy).

Provider Input Output Cache read /M Latency Throughput Uptime
Google ~$0.012/min ~$0.012/min ~150ms ~150 min/s 100.00%
Azure ~$0.013/min ~$0.013/min ~170ms ~120 min/s 100.00%
Amazon Web Services ~$0.014/min ~$0.014/min ~190ms ~100 min/s 100.00%

Try this model

Test Chirp 3 right here — free to start.

Chirp 3
Hi! Want to test the model?

Suggestions for your first prompt

Code snippet

Call the model through the OpenAI-compatible API.

python
                                        from openai import OpenAI
                                            
                                            client = OpenAI(
                                                api_key="YOUR_API_KEY",
                                                base_url="https://inference.example.com/v1"
                                            )
                                            
                                            response = client.chat.completions.create(
                                                model="google/chirp-3",
                                                messages=[
                                                    {
                                                        "role": "user",
                                                        "content": "Describe this image in one sentence."
                                                    }
                                                ],
                                            )
                                            
                                            print(response.to_json())
                                        
                                    
                                        {
                                                "model": "google/chirp-3",
                                                "messages": [
                                                    {
                                                        "role": "user",
                                                        "content": "Describe this image in one sentence."
                                                    }
                                                ]
                                            }
                                        
                                    

Uptime

30-Day Uptime
100.00%
Past Incidents (30d)
0
Error rate (24h)
0.00%

Last 30 days

30/30 days operational | 100.00% uptime

30 days ago Today
Operational Degraded Outage Maintenance
See All Incidents

5 Core Capabilities

  • Conversational AI

    Engages in natural, multi-turn voice conversations, understanding user intent and context to provide relevant, coherent spoken responses.

  • Audio Transcription

    Converts spoken language in audio input into accurate text, supporting real-time or near real-time voice transcription scenarios.

  • Speech Translation

    Translates spoken language from one language to another, enabling cross-lingual voice conversations and real-time interpretation use cases.

  • Voice Monitoring

    Processes and monitors audio streams for commands or triggers, enabling responsive voice-driven applications and interactive systems.

  • Audio-Linked Imagery

    Can be integrated with image-capable systems to associate spoken descriptions with visual content, supporting multimodal user experiences.

6 Most Valuable Use Cases

  • Real-time transcription
  • Call center analytics
  • Meeting note generation
  • Customer support voicebots
  • Audiobook narration
  • Custom brand voices

Why Build on LLM.API?

One unified API. Every major model. Built-in reliability, cost control, and observability.

  • Unified AI Routing

    Automatically route each request to the optimal model across providers based on latency, cost, and quality. One endpoint, dynamic policies, no SDK sprawl.

    One endpoint, any model
  • Cost-Aware Orchestration

    Control spend with price-aware routing, per-project limits, and transparent metering across vendors. Swap models without rewiring billing or touching client code.

    Cut cost, keep quality
  • Resilient Fallback Flows

    Design multi-provider fallback trees that auto-retry on failures, timeouts, or quota limits. Keep production workloads online even when a vendor has issues.

    Never ship single-vendor SPOF
  • Deep LLM Observability

    Get unified traces, logs, and metrics for every request across providers. Inspect prompts, latencies, and errors in one place to debug faster and tune confidently.

    Single pane of AI truth
  • Task-Level Abstractions

    Describe tasks like “chat”, “embed”, or “moderate” instead of binding to model names. LLM.API maps tasks to the best capabilities behind a stable interface.

    Code to tasks, not models
  • High-Throughput Batch APIs

    Ship bulk workloads with streaming-safe, rate-aware batching. Push thousands of prompts per job while LLM.API handles chunking, retries, and provider limits.

    Batch at production scale

When to Use — When NOT to Use

Use it if...

  • You need a general-purpose chat model for consumer-facing assistants or help bots.
  • Your use case involves everyday Q&A, explanations, and basic task automation workflows.
  • You need tight integration with Google ecosystems, tooling, or existing Google Cloud infrastructure.
  • Your use case involves moderate-length documents where natural language understanding is more important than depth.
  • You need reasonably capable text generation without requiring cutting-edge reasoning or niche domain expertise.
  • Your use case involves prototyping conversational features before committing to a more advanced model.

Avoid if...

  • You need state-of-the-art complex reasoning, planning, or tool-using capabilities across long sessions.
  • Your workload requires rigorous handling of long technical documents with precise, verifiable citations.
  • You need highly optimized performance on specialized domains like law, medicine, or quantitative finance.
  • Your workload requires extremely long context windows with consistent accuracy across hundreds of pages.
  • You need fine-grained control over safety, customization, or model behavior beyond standard configuration options.
  • Your workload requires strict reproducibility and deterministic outputs for compliance-critical pipelines.

Frequently Asked Questions

  • What is Chirp 3?

    Chirp 3 is a Google speech model focused on automatic speech recognition with strong multilingual performance and robustness to noisy, real‑world audio.

  • What is Chirp 3 best suited for?

    Chirp 3 is best for high‑accuracy, large‑scale transcription of calls, meetings, videos, and user‑generated audio across many languages and accents.

  • What modalities does Chirp 3 support through LLM.API?

    Through LLM.API, Chirp 3 supports audio input and text output for speech‑to‑text workloads, without image or text‑generation capabilities.

  • How is Chirp 3 priced on LLM.API?

    Chirp 3 is typically billed per processed audio minute or second via LLM.API; check your LLM.API pricing page for exact current rates.

  • What is the maximum audio or context length Chirp 3 can handle?

    Chirp 3 supports long‑form audio transcription, but maximum duration and effective context depend on LLM.API limits and configuration for streaming or batch mode.

  • How fast is Chirp 3 in terms of latency?

    Chirp 3 generally operates near real time for short clips, with latency mainly determined by audio length and LLM.API region and network conditions.

  • How do I call Chirp 3 via the LLM.API?

    You select the Google Chirp 3 model in your LLM.API request, provide audio bytes or a URL, and receive transcribed text in the response.

  • How does Chirp 3 compare to general LLMs for transcription tasks?

    Compared with general text LLMs, Chirp 3 is specialized, usually cheaper and more accurate for speech recognition but cannot perform text‑only reasoning.

  • Does Chirp 3 support streaming transcription on LLM.API?

    If enabled by LLM.API, Chirp 3 can consume audio chunks incrementally and return partial transcripts for low‑latency streaming experiences.

  • What are the main limitations of Chirp 3?

    Chirp 3 is limited to speech recognition, may struggle with extremely noisy audio, rare languages, domain‑specific jargon, and does not generate or understand images.

Get one key to every model

Swap your API key. Keep your code.