Wan 2.7
Wan 2.7 is Alibaba’s latest open-source multimodal visual generation model for high-quality video and image creation, offering text-to-video, image-to-video, text-to-image, and editing in a single architecture.
What is Wan 2.7?
Wan 2.7 is an AI visual generation model from Alibaba that unifies video and image generation and editing in one system. It is mainly used for generating cinematic short videos from text or image prompts and for image-to-video transformations in creative, marketing, and storytelling workflows. It is also used for text-to-image creation, reference-guided image generation, and instruction-based image or video editing for design and content production teams. Wan 2.7 is part of Alibaba’s Wan video model family developed within the broader Qwen ecosystem, succeeding earlier Wan 2.x releases.
Providers
Different companies host the same model. OpenRouter routes your request to one of them based on the routing mode you pick — Balanced (price + speed), Nitro (fastest), or Exacto (highest tool-calling accuracy).
| Provider | Input | Output | Cache read /M | Latency | Throughput | Uptime |
|---|---|---|---|---|---|---|
| Alibaba Cloud | ~$0.40 | ~$1.20 | — | ~260ms | ~60 tps | ~99.95% |
| OpenAI | ~$0.50 | ~$1.50 | — | ~180ms | ~90 tps | 99.06% |
| Azure AI | ~$0.55 | ~$1.60 | — | ~200ms | ~80 tps | 100.00% |
| Anthropic | ~$0.60 | ~$1.80 | — | ~190ms | ~70 tps | 99.28% |
Try this model
Test Wan 2.7 right here — free to start.
Suggestions for your first prompt
Code snippet
Call the model through the OpenAI-compatible API.
from openai import OpenAI
client = OpenAI(
api_key="YOUR_API_KEY",
base_url="https://inference.example.com/v1"
)
response = client.chat.completions.create(
model="alibaba/wan-2-7",
messages=[
{
"role": "user",
"content": "Describe this image in one sentence."
}
],
)
print(response.to_json())
{
"model": "alibaba/wan-2-7",
"messages": [
{
"role": "user",
"content": "Describe this image in one sentence."
}
]
}
5 Core Capabilities
-
Text-to-Video
Generates high-quality video clips directly from detailed text prompts, supporting controllable camera movement, scenes, and lighting for creators.
-
Image-to-Video
Animates still images into coherent motion videos, preserving subject appearance and layout while adding realistic movement and transitions.
-
Reference-Based Editing
Edits and extends existing video using reference frames and instructions, enabling consistent subjects, motion control, and frame-level refinements.
-
Unified Visual Suite
Acts as a multimodal visual model handling text-to-video, image-to-video, text-to-image, and image editing within a single architecture.
-
Thinking Mode Control
Interprets user intent before rendering with a dedicated thinking phase, improving creative consistency, controllability, and reducing failed generations.
6 Most Valuable Use Cases
- Text-to-Image Generation
- Image-to-Video Animation
- Text-to-Video Creation
- Video-to-Video Editing
- Marketing Visual Production
- Multimodal Media Research
Why Build on LLM.API?
One unified API. Every major model. Built-in reliability, cost control, and observability.
-
Unified AI Routing
Dynamically route each request to the optimal model across providers based on performance, cost, and availability—without changing your code or client integration.
One endpoint, any model -
Cost-Aware Orchestration
Automatically balance premium and budget models with configurable cost ceilings, so you keep latency low and quality high while tightly controlling spend.
Optimize quality per dollar -
Resilient Fallback Logic
Define provider and model failover chains so requests transparently retry on alternates, insulating your app from regional outages, rate limits, or model regressions.
Stay online, even upstream -
Deep Model Observability
Get unified traces, latency, error, and token metrics across all providers with request-level logs for fast debugging, tuning, and capacity planning.
See every token, everywhere -
Task-Level Abstractions
Call high-level tasks—chat, generation, tools, and more—instead of vendor-specific APIs, so you can swap models without rewriting application logic.
Program to tasks, not models -
High-Throughput Batch APIs
Submit large batches of prompts in a single call with automatic chunking, concurrency control, and retries to maximize throughput and minimize overhead.
Scale workloads, not code
When to Use — When NOT to Use
Use it if...
- You need a strong general-purpose Chinese language model from a major Chinese provider.
- You need reasonably capable text generation for chatbots, assistants, or content drafting.
- You need to integrate with Alibaba Cloud services or an existing Alibaba ecosystem.
- Your use case involves moderate reasoning tasks that do not require frontier-level performance.
- Your use case involves experimentation with multiple Chinese LLMs, including non–US-based offerings.
- You need a vendor-diverse backup model where Western foundation models are restricted.
Avoid if...
- You need state-of-the-art reasoning, coding, or tool-use comparable to the latest frontier models.
- Your workload requires detailed, up-to-date knowledge of non-Chinese global regulatory landscapes.
- You need guaranteed support for highly specialized domains like advanced biotech, aerospace, or cryptography.
- You need the broadest ecosystem of third-party tools, plugins, and community examples available.
- Your workload requires clear, well-documented compliance attestations for US or EU-specific regulations.
- You need extremely transparent, English-first documentation and debugging resources for all model behaviors.
Frequently Asked Questions
-
What is Wan 2.7?
Wan 2.7 is an Alibaba large language model accessible via LLM.API, targeting general-purpose text generation and understanding tasks.
-
What is Wan 2.7 best suited for?
Wan 2.7 is best for cost-efficient chatbots, content generation, and general NLP tasks where balanced quality and efficiency matter.
-
What is the context window of Wan 2.7?
Wan 2.7 supports a context window of up to 8,192 tokens via LLM.API.
-
How fast is Wan 2.7 on LLM.API?
Wan 2.7 is optimized for low latency on LLM.API, typically returning first tokens within a few hundred milliseconds under normal load.
-
Which modalities does Wan 2.7 support?
Wan 2.7 is a text-only model on LLM.API, supporting text inputs and text outputs.
-
How is Wan 2.7 priced on LLM.API?
Wan 2.7 uses LLM.API’s unified token-based billing, with separate input and output token rates shown in your LLM.API pricing dashboard.
-
How do I call Wan 2.7 through LLM.API?
You select provider 'Alibaba' and model 'Wan 2.7' in the LLM.API request payload, keeping the standard chat or completion schema unchanged.
-
How does Wan 2.7 compare to similar models on LLM.API?
Wan 2.7 generally trades slightly lower peak quality than top-tier frontier models for better cost efficiency and predictable performance.
-
Does Wan 2.7 support streaming responses on LLM.API?
Yes, Wan 2.7 supports token streaming via LLM.API by enabling the standard 'stream' flag in your request.
-
What are key limitations of Wan 2.7?
Wan 2.7 may struggle with highly specialized domain knowledge, strict mathematical reasoning, and tasks requiring very long-context retention beyond its context window.
COMPARE
Competitive Models
-
Claude Opus 4.6
Claude Opus 4.6 is a large language model from Anthropic’s Claude Opus series, designed as a high-end, general-purpose AI assistant with strong reasoning and language capabilities. It is notable for being one of Anthropic’s flagship frontier models, aimed at complex tasks requiring advanced comprehension and generation.
-
GLM 4.6V
GLM 4.6V is Z.ai’s open-source, large-scale vision-language model that supports images, video, documents, and text with a long context window and native tool use. It is notable for combining high-quality multimodal understanding with function calling and cloud- or local-friendly variants.
-
MiniMax M2
MiniMax M2 is an open‑weight Mixture‑of‑Experts large language model from MiniMax, designed to deliver high coding and agentic workflow performance with low latency and cost. It uses 230B total parameters with only about 10B active per token to balance strong reasoning with efficient deployment.
Get one key to every model
Swap your API key. Keep your code.