提供用于聊天和编码的 Qwen 等强大开源与专有模型,以及用于视频和图像生成的 Wan 2.x 系列。

Alibaba 的模型生态覆盖 Qwen 对话模型和 Wan 视觉生成系列。MuAPI 将 Alibaba 模型与其他提供商并列,帮助团队比较能力、价格和适合的集成方式。

Alibaba AI models on MuAPI
返回提供商目录
探索/Alibaba 模型
Alibaba

Alibaba AI API 模型

MuAPI 上的 Alibaba AI 模型

探索 Alibaba 的聊天、代码、图像和视频生成模型,包括可通过 MuAPI 使用的 Gemini、Nano Banana 和 Veo 类工作流。

全部模型

90 个模型

视频生成模型

视频

$0.700 /

happy-horse-1.1-text-to-video-720p

Happy Horse 1.1 Text to Video (720p) — generate expressive 720p video clips from text prompts with vivid character motion.

wan3.0-text-to-video
视频

$0.500 /

wan3.0-text-to-video

Wan 3.0 Text to Video generates a video with synchronized audio directly from a text prompt, with selectable resolution, aspect ratio, and duration.

wan3.0-prime-reference-to-video
视频

$0.700 /

wan3.0-prime-reference-to-video

Wan 3.0 Prime Reference to Video generates a higher-fidelity video guided by a prompt plus up to 10 reference images, 5 reference videos, and 5 reference audios for visual and audio coherence.

视频

$0.300 /

wan2.2-image-to-video

Wan 2.2’s I2V mode brings static visuals to life with vivid, expressive animations. It interprets motion, emotion, and background dynamics from a single image to generate smooth and cinematic short videos.

视频

$0.300 /

wan2.1-text-to-video

WAN 2.1 turns your written prompts into vivid, cinematic video clips. Ideal for storytelling, content creation, and visualizing abstract ideas, it supports detailed natural scenes, character motion, and dramatic camera movements — all from just text.

视频

$0.350 /

wan2.2-animate

Wan2.2 Animate is a video-to-video model for animating a character or replacing a character in existing video clips. It replicates holistic movement and facial expressions from a reference video or pose while preserving the target character’s appearance. You upload both an image (for the character) and a video containing motion/expression, and the model generates a video where the character in your image moves like the reference. Supports 480p or 720p, up to 120 seconds

视频

$0.300 /

wan2.2-text-to-video

Wan 2.2’s T2V mode transforms descriptive text prompts into high-quality, stylized video sequences. It excels at generating anime-style or cinematic visuals with smooth motion and strong thematic consistency.

wan3.0-spicy-text-to-video
视频

$0.550 /

wan3.0-spicy-text-to-video

Wan 3.0 Spicy Text to Video generates a bold, high-contrast, high-motion video with synchronized audio directly from a text prompt, with selectable resolution, aspect ratio, and duration.

视频

$0.100 /

wan2.1-reference-video

WAN 2.1 is an advanced AI model that transforms one or more reference images into a coherent, animated video. By combining characters, objects, or environments from multiple images, it creates smooth motion sequences while preserving realism, style, and fine details.

视频

$0.300 /

wan2.2-edit-video

Easily modify existing videos using simple text commands. With Wan 2.2 Video-Edit, you can change attire, character appearance, or other visual elements directly within your video—no need to start from scratch. Works on uploads of 480p or 720p, for up to two minutes.

视频

$0.650 /

wan2.6-image-to-video

WAN 2.6 Image-to-Video converts a single still image into a smooth, cinematic video clip. It preserves the original image’s composition, lighting, and style while adding natural motion, depth parallax, atmospheric effects, and gentle camera movement.

视频

$0.200 /

wan2.2-spicy-image-to-video

Wan2.2-spicy Image-to-Video transforms a single creative image into a short dynamic video with bold motion, stylized effects, high-contrast lighting, and energy-driven animations. The “spicy” variant produces more dramatic movement, more vivid colors, and more expressive visual effects.

视频

$1.050 /

happy-horse-1-video-edit-720p

Happy Horse 1.0 Video Edit (720p) - modify an input video at 720p using a natural-language instruction with optional reference images.

wan2.7-text-to-video
视频

$0.100 /

wan2.7-text-to-video

Alibaba WAN 2.7 Text-to-Video turns plain prompts into coherent, cinematic clips.

wan2.7-image-to-video
视频

$0.100 /

wan2.7-image-to-video

Alibaba WAN 2.7 converts images into videos with optional audio.

wan2.7-reference-to-video
视频

$0.100 /

wan2.7-reference-to-video

Alibaba WAN 2.7 Reference-to-Video. Reference characters/props to generate new shots.

wan2.7-video-extend
视频

$0.100 /

wan2.7-video-extend

Extend existing videos seamlessly with Wan 2.7.

wan2.7-video-edit
视频

$0.100 /

wan2.7-video-edit

Perform prompt-driven video editing with multi-image reference support.

视频

$1.800 /

happy-horse-1-text-to-video-1080p

Happy Horse 1.0 Text to Video — generate expressive, stylized video clips from text prompts with vivid character motion and dynamic scene storytelling.

视频

$1.800 /

happy-horse-1-image-to-video-1080p

Happy Horse 1.0 Image to Video — bring still images to life with fluid, expressive animation and fine-grained motion control.

视频

$0.900 /

happy-horse-1-image-to-video-720p

Happy Horse 1.0 Image to Video (720p) — bring still images to life with fluid, expressive animation at 720p output resolution.

视频

$0.900 /

happy-horse-1-text-to-video-720p

Happy Horse 1.0 Text to Video (720p) — generate expressive, stylized video clips from text prompts at 720p output resolution.

视频

$2.100 /

happy-horse-1-reference-to-video-1080p

Happy Horse 1.0 Reference to Video (1080p) - generate expressive 1080p video clips conditioned on 1-9 reference images plus a text prompt.

视频

$1.050 /

happy-horse-1-reference-to-video-720p

Happy Horse 1.0 Reference to Video (720p) - generate expressive 720p video clips conditioned on 1-9 reference images plus a text prompt.

视频

$0.900 /

happy-horse-1.1-text-to-video-1080p

Happy Horse 1.1 Text to Video (1080p) — generate expressive 1080p video clips from text prompts with vivid character motion and dynamic scene storytelling.

视频

$0.900 /

happy-horse-1.1-image-to-video-1080p

Happy Horse 1.1 Image to Video (1080p) — bring still images to life with fluid, expressive 1080p animation.

视频

$0.700 /

happy-horse-1.1-image-to-video-720p

Happy Horse 1.1 Image to Video (720p) — bring still images to life with fluid, expressive 720p animation.

视频

$0.700 /

happy-horse-1.1-reference-to-video-720p

Happy Horse 1.1 Reference to Video (720p) — generate 720p video conditioned on 1-9 reference images plus a text prompt.

视频

$0.900 /

happy-horse-1.1-reference-to-video-1080p

Happy Horse 1.1 Reference to Video (1080p) — generate 1080p video conditioned on 1-9 reference images plus a text prompt.

视频

$0.900 /

happy-horse-1.1-video-edit-1080p

Happy Horse 1.1 Video Edit (1080p) — modify an input video using natural-language instructions with optional reference images.

视频

$0.700 /

happy-horse-1.1-video-edit-720p

Happy Horse 1.1 Video Edit (720p) — modify an input video using natural-language instructions with optional reference images.

视频

$0.016 /

wan2.2-5b-fast-t2v

Wan 2.2 Fast is a lightweight, high-speed version of the Wan 2.2 model, optimized for quick text-to-video generation. It trades some cinematic detail for rapid results, making it perfect for prototyping, previews, social media clips, and quick storytelling.

视频

$0.200 /

wan2.2-speech-to-video

WAN2.2 Speech-to-Video transforms a static image into a talking video by synchronizing lip movements and facial expressions with an audio input. Simply provide a character image along with a speech dialogue, and the model generates a natural, expressive video where the subject speaks your lines.

视频

$0.300 /

wan2.1-image-to-video

Animate static images into expressive video sequences with WAN 2.1. Upload any image and guide its transformation into a moving scene — great for bringing art, characters, or photos to life with smooth motion and consistent style.

视频

$0.650 /

wan2.5-image-to-video

WAN 2.5 Image-to-Video takes your image as the starting frame and turns it into a dynamic video, preserving realism, motion, and camera effects. Upload a static image, add a descriptive text prompt, and the model generates cinematic motion—camera pans, environmental movement, and realistic physics—across the result.

视频

$0.650 /

wan2.5-text-to-video

WAN 2.5 Text-to-Video transforms written prompts into cinematic video clips with dynamic motion, realistic physics, and natural animation. It can also generate characters delivering dialogue, making it ideal for storytelling, ads, and creative showcases.

视频

$0.440 /

wan2.5-image-to-video-fast

Convert a single static image into a cinematic short video with realistic motion, dynamic camera movement, and environmental effects. The Fast mode generates high-quality videos quickly, perfect for rapid prototyping, social media clips, and immersive visual storytelling from still images.

视频

$0.440 /

wan2.5-text-to-video-fast

Transform text prompts into short, cinematic videos with natural motion, realistic environments, and dynamic camera perspectives. Fast mode delivers quick, high-fidelity video generation, ideal for creative storytelling, concept visuals, and social media content.

wan3.0-image-to-video
视频

$0.500 /

wan3.0-image-to-video

Wan 3.0 Image to Video animates a source image with synchronized audio from a motion prompt, with an optional end-frame image, selectable resolution, aspect ratio, and duration.

视频

$0.200 /

wan2.2-spicy-video-extend

Wan-2.2-spicy Video Extend continues an existing video by generating new frames that match the original style but add stronger motion, bolder effects, and spicier dramatics.

视频

$0.650 /

wan2.6-text-to-video

WAN 2.6 Text-to-Video generates smooth, cinematic videos directly from text prompts. It’s designed for strong scene coherence, atmospheric depth, and fluid camera motion, making it ideal for fantasy and sci-fi worlds, surreal concepts, environmental storytelling, and dramatic visual sequences with rich lighting and motion.

wan3.0-spicy-image-to-video
视频

$0.550 /

wan3.0-spicy-image-to-video

Wan 3.0 Spicy Image to Video animates a source image into a bold, high-contrast, high-motion video with synchronized audio from a motion prompt, with an optional end-frame image, selectable resolution, aspect ratio, and duration.

视频

$2.100 /

happy-horse-1-video-edit-1080p

Happy Horse 1.0 Video Edit (1080p) - modify an input video at 1080p using a natural-language instruction with optional reference images.

wan2.6-image-to-video-spicy
视频

$1.000 /

wan2.6-image-to-video-spicy

Alibaba WAN 2.6 Spicy Image-to-Video animates a still image into a bold, high-motion clip with optional audio-guided generation, single- or multi-shot framing, and prompt-expansion.

wan2.7-image-to-video-spicy
视频

$1.000 /

wan2.7-image-to-video-spicy

Alibaba WAN 2.7 Spicy Image-to-Video animates a still image into a bold, high-motion clip with optional audio-guided generation, single- or multi-shot framing, and prompt-expansion.

wan3.0-prime-text-to-video
视频

$0.700 /

wan3.0-prime-text-to-video

Wan 3.0 Prime Text to Video generates a higher-fidelity video with synchronized audio directly from a text prompt, with selectable resolution, aspect ratio, and duration.

wan3.0-reference-to-video
视频

$0.500 /

wan3.0-reference-to-video

Wan 3.0 Reference to Video generates a video guided by a prompt plus up to 10 reference images, 5 reference videos, and 5 reference audios for visual and audio coherence.

wan3.0-prime-image-to-video
视频

$0.700 /

wan3.0-prime-image-to-video

Wan 3.0 Prime Image to Video animates a source image with synchronized audio from a motion prompt at higher fidelity, with an optional end-frame image, selectable resolution, aspect ratio, and duration.

wan3.0-spicy-reference-to-video
视频

$0.550 /

wan3.0-spicy-reference-to-video

Wan 3.0 Spicy Reference to Video generates a bold, high-contrast, high-motion video guided by a prompt plus up to 10 reference images, 5 reference videos, and 5 reference audios for visual and audio coherence.

图像生成模型

z-image-base
图像

$0.013 / 1K tokens

z-image-base

Z-Image Base is a general-purpose text-to-image model designed for reliable, high-quality image generation from natural language prompts. It focuses on clear composition, good prompt adherence, and versatile output across everyday scenes, product-style visuals, characters, and creative concepts.

qwen-text-to-image-2512
图像

$0.040 / 次生成

qwen-text-to-image-2512

Qwen Image Text-to-Image 2512 generates high-resolution, visually consistent images from text prompts. It focuses on strong scene structure, clean composition, and atmospheric lighting, making it well-suited for cinematic environments, surreal concepts, fantasy and sci-fi worlds.

wan2.1-text-to-image
图像

$0.030 / 1K tokens

wan2.1-text-to-image

WAN 2.1 is a powerful AI model that transforms text prompts into high-resolution, photorealistic images. It excels at detailed object rendering, realistic lighting, and fine textures, making it ideal for visual content, concept art, advertising, and digital storytelling.

wan2.7-text-to-image
图像

$0.050 / 1K tokens

wan2.7-text-to-image

Alibaba WAN 2.7 Text-to-Image generates high-quality images from text prompts with thinking mode for enhanced image quality.

图像

$0.300 / 次生成

wan2.1-lora-i2v

Bring still images to life using WAN 2.1 LoRA I2V, which supports custom LoRA fine-tunes for identity consistency. Animate expressions, subtle movements, or full-body actions while preserving personalized features from the image and LoRA.

图像

$0.300 / 次生成

wan2.1-lora-t2v

WAN 2.1 LoRA T2V enables users to generate videos from text prompts with custom-trained LoRA modules. Tailor the generation to specific characters, outfits, or animation styles — ideal for brand storytelling, fan content, and stylized animations.

qwen-image
图像

$0.030 / 1K tokens

qwen-image

Generate high-quality, detailed images from text prompts in various styles — from realistic to artistic — perfect for creative visuals, product shots, and concept art.

qwen-image-edit-plus
图像

$0.030 / 次生成

qwen-image-edit-plus

Qwen Image Edit Plus is an upgraded image-editing model that supports multiple image references and superior text editing. Powered by the 20B-parameter Qwen architecture, it allows changes like background swap, style transfer, object removal/addition, and precise text edits (bilingual: English/Chinese) while maintaining visual consistency and preserving details of the original images.

z-image-lora-trainer
图像

$2.500 / 次生成

z-image-lora-trainer

Train custom image LoRA models from your dataset with zip uploads, auto-tuned defaults, and fast iteration for brand, character, or IP looks.

qwen-image-2512-lora-trainer
图像

$2.000 / 次生成

qwen-image-2512-lora-trainer

Train custom Qwen-Image-2512 LoRA models 10x faster with style, character, and object training. Upload a ZIP dataset with images to start.

wan2.5-text-to-image
图像

$0.040 / 1K tokens

wan2.5-text-to-image

WAN 2.5 Text-to-Image generates high-quality, realistic or stylized images from textual descriptions. It supports detailed visual storytelling, cinematic compositions, and versatile styles — from portraits and product shots to landscapes and fantasy scenes.

qwen-image-edit-2511
图像

$0.040 / 次生成

qwen-image-edit-2511

Qwen Image Edit 2511 performs precise, instruction-driven edits on an existing image while preserving composition, lighting, and overall style. It’s well-suited for object replacement, material changes, localized edits, and subtle scene adjustments with strong visual consistency and minimal artifacts.

wan2.6-image-edit
图像

$0.045 / 次生成

wan2.6-image-edit

WAN 2.6 Image Edit applies targeted, instruction-based edits to an existing image while preserving composition, perspective, and lighting. It’s ideal for object replacement, material changes, environment tweaks, and style adjustments with clean integration and minimal artifacts—keeping the original scene coherent and cinematic.

z-image-turbo
图像

$0.007 / 1K tokens

z-image-turbo

Z-Image Turbo is a high-speed text-to-image model optimized for fast creative generation. It produces detailed, high-contrast, high-resolution images with strong stylization control. Ideal for rapid concept creation, visual exploration, product ideas, fantasy scenes, and cinematic composition tests. Designed for low latency and strong prompt adherence.

z-image-p
图像

$0.004 / 1K tokens

z-image-p

Z-Image P is based on PiAPI's Qubico/z-image text-to-image model.

wan2.7-image-edit
图像

$0.050 / 次生成

wan2.7-image-edit

Alibaba WAN 2.7 Image Edit performs prompt-driven image editing with support for multiple-image references.

qwen-image-2.0
图像

$0.040 / 1K tokens

qwen-image-2.0

Qwen 2.0 Text to Image model with enhanced realism.

qwen-image-2.0-pro
图像

$0.090 / 1K tokens

qwen-image-2.0-pro

Qwen 2.0 Pro Text to Image model with maximum realism and fidelity.

qwen-image-2.0-pro-edit
图像

$0.090 / 次生成

qwen-image-2.0-pro-edit

Qwen 2.0 Pro Image Edit model with maximum precision and modifications.

wan2.7-text-to-image-pro
图像

$0.100 / 1K tokens

wan2.7-text-to-image-pro

Alibaba WAN 2.7 Text-to-Image Pro generates high-quality images up to 4K from text prompts with thinking mode for enhanced image quality.

wan3.0-image-edit
图像

$0.050 / 次生成

wan3.0-image-edit

Wan 3.0 Image Edit is an upcoming AI image-editing model. Confirmed API controls, output specifications, and pricing will be published at launch.

qwen-image-edit
图像

$0.030 / 次生成

qwen-image-edit

The Qwen Edit Image Model allows you to modify existing images using text-based editing prompts. Instead of generating from scratch, you can upload a base image and describe the desired changes (e.g., replacing objects, altering colors, adding new elements).

wan3.0-text-to-image
图像

$0.050 / 1K tokens

wan3.0-text-to-image

Wan 3.0 Text to Image is an upcoming AI image model. Confirmed API controls, output specifications, and pricing will be published at launch.

wan2.5-image-edit
图像

$0.040 / 次生成

wan2.5-image-edit

The Wan2.5 Edit Image model allows you to transform existing images with precision and creativity. By providing an image along with an edit prompt, you can make realistic changes, enhancements, or stylistic adjustments—whether it’s altering objects, changing backgrounds, adding details, or applying an entirely new artistic style.

qwen3-pro-image-to-image
图像

$0.040 / 次生成

qwen3-pro-image-to-image

Qwen 3.0 Pro Image to Image transforms and edits existing reference images according to text prompts with professional-grade precision.

qwen3-pro-text-to-image
图像

$0.040 / 1K tokens

qwen3-pro-text-to-image

Qwen 3.0 Pro Text to Image generates high-fidelity professional photorealistic images and editorial artwork from prompts with enhanced detail.

qwen-image-lora-trainer
图像

$2.000 / 次生成

qwen-image-lora-trainer

Train custom Qwen-Image LoRA models 10x faster. Fine-tune styles, characters, or object concepts from a ZIP dataset with auto-tuned learning rates and steps.

z-image-base-lora-trainer
图像

$2.500 / 次生成

z-image-base-lora-trainer

Train custom Z-Image Base LoRA models from your dataset with zip uploads, auto-tuned defaults, and fast iteration for brand, character, or IP looks.

wan2.6-text-to-image
图像

$0.040 / 1K tokens

wan2.6-text-to-image

WAN 2.6 Text-to-Image generates detailed, cinematic still images from text prompts. It focuses on strong composition, atmospheric lighting, and clear subject structure, making it suitable for fantasy and sci-fi environments, surreal concepts, architectural visuals, and dramatic world-building imagery.

qwen3-text-to-image
图像

$0.030 / 1K tokens

qwen3-text-to-image

Qwen 3.0 Text to Image generates high-fidelity photorealistic images and editorial artwork from prompts with intelligent prompt rewriting.

qwen3-image-to-image
图像

$0.030 / 次生成

qwen3-image-to-image

Qwen 3.0 Image to Image transforms and edits existing reference images according to text prompts with precise style and content preservation.

qwen-image-2.0-edit
图像

$0.040 / 次生成

qwen-image-2.0-edit

Qwen 2.0 Image Edit model with precise background modification and enhancements.

wan2.7-image-edit-pro
图像

$0.100 / 次生成

wan2.7-image-edit-pro

Alibaba WAN 2.7 Image Edit Pro performs prompt-driven image editing with multi-image reference support and up to 2K output.

其他工具模型

qwen-image-text-to-image-lora
Lora Support

$0.020 / 次生成

qwen-image-text-to-image-lora

Qwen-Image Text-to-Image LoRA pairs the 20B MMDiT next-generation text-to-image model with custom LoRA adapters (up to 3 LoRAs) for fast style customization, refined aesthetics, and character consistency.

qwen-image-edit-lora
Lora Support

$0.040 / 次生成

qwen-image-edit-lora

Qwen Image Edit LoRA is a fine-tuned image editing model with custom LoRA support. Modifying or adding elements with fine-tuned styles.

qwen-image-edit-plus-lora
Lora Support

$0.040 / 次生成

qwen-image-edit-plus-lora

Qwen-Image-Edit-Plus (2509) is 20B MMDiT image-to-image editor supporting multi-image edits, single-image consistency, and native ControlNet. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

z-image-base-text-to-image-lora
Lora Support

$0.010 / 次生成

z-image-base-text-to-image-lora

Z-Image-Base LoRA (6B) text-to-image and image-to-image generation model with external LoRA support (up to 3 LoRAs) and aspect ratio selection.

z-image-turbo-text-to-image-lora
Lora Support

$0.010 / 次生成

z-image-turbo-text-to-image-lora

Z-Image-Turbo LoRA (6B) enables ultra-fast text-to-image generation with external LoRA support. Generate photorealistic images in sub-second latency while applying up to 3 LoRAs for custom styles.

qwen-image-text-to-image-2512-lora
Lora Support

$0.020 / 次生成

qwen-image-text-to-image-2512-lora

Qwen-Image Text-to-Image 2512 LoRA pairs the 20B MMDiT next-generation text-to-image model with custom LoRA adapters (up to 3 LoRAs) for fast style customization, refined aesthetics, and character consistency.

z-image-turbo-image-to-image-lora
Lora Support

$0.010 / 次生成

z-image-turbo-image-to-image-lora

Z-Image-Turbo Image-to-Image LoRA transforms reference images with custom LoRA styles in sub-second time. Apply up to 3 LoRAs for personalized image transformation.

qwen-image-edit-2511-lora
Lora Support

$0.040 / 次生成

qwen-image-edit-2511-lora

Qwen Image Edit 2511 LoRA is an enhanced version with custom LoRA support for personalized styles. It delivers stronger edit consistency, robust multi-person identity/pose consistency, custom LoRA styles, enhanced industrial/product design, and improved geometric reasoning for structure-preserving edits.