Ovi 是统一模型,可从文本入力生成同步的视频与音频。你可以编写包含对白和环境声音的场景描述,Ovi 会生成一段短视频(通常约 5 秒),让画面与声音自然对齐。视频以 540p 分辨率生成。
このモデルについて
Ovi 是一款前沿的统一模型,能够根据文本入力无缝生成同步的视频和音频。用户只需一个プロンプト,就能创建短小且引人入胜的视频片段,自然融合画面与声音,非常适合数字叙事、广告和快速多媒体原型制作。Ovi 基于复杂的生成算法,确保每段视频不仅视觉吸引人,声音也能保持同步,在清晰的 540p 分辨率下呈现近似传统拍摄的体验。
该模型利用先进的机器学习技术理解详细的场景描述、对白和环境声音,并将它们映射为连贯的视频出力。这让创作者可以轻松将想象转化为有形媒体,弥合文本概念与动态视听内容之间的差距。其简洁的界面和高效处理能力使其成为极具竞争力且经济高效的解决方案,让用户只需极少入力和较短周转时间即可获得高质量结果。
コスト分析
| プロバイダー | 費用 | 備考 |
|---|---|---|
| muapiapp | $0.20 per generation | muapiapp is 20-50% more affordable compared to competitors while delivering comparable or superior quality. |
| Fal.ai | $0.25 per generation | muapiapp offers a 20% discount over Fal.ai, providing excellent cost efficiency with high-quality video and audio outputs. |
| Replicate | $0.25 per generation | muapiapp is priced at 20% less than Replicate, ensuring a more cost-effective solution without compromising on performance. |
muapiapp is 20-50% more affordable compared to competitors while delivering comparable or superior quality.
muapiapp offers a 20% discount over Fal.ai, providing excellent cost efficiency with high-quality video and audio outputs.
muapiapp is priced at 20% less than Replicate, ensuring a more cost-effective solution without compromising on performance.
** 競合サービスの料金は類似のモデル構成および利用ティアに基づいて算出された推定値です。
設定スキーマ
| パラメータ | 型 | 説明 | デフォルト |
|---|---|---|---|
| プロンプト | string | 描述動画内容的文本プロンプト。 | A busy railway station intersection at day time. A loudspeaker voice announces: <S>Attention: Train 12785 to Mumbai is now boarding at Platform 3.<E>. Ambient: <AUDCAP>Car horns, crowd murmur, distant train whistle<ENDAUDCAP>. |
| アスペクト比 | Enum(2個の選択肢) | 出力動画的アスペクト比。 | 16:9 |
描述動画内容的文本プロンプト。
A busy railway station intersection at day time. A loudspeaker voice announces: <S>Attention: Train 12785 to Mumbai is now boarding at Platform 3.<E>. Ambient: <AUDCAP>Car horns, crowd murmur, distant train whistle<ENDAUDCAP>.出力動画的アスペクト比。
16:9開発者ドキュメント
使用方法 Ovi-text-to-video
準備入力:
<S> 表示对白,<AUDCAP> 表示环境音频)。A busy railway station intersection at day time. A loudspeaker voice announces: <S>Attention: Train 12785 to Mumbai is now boarding at Platform 3.<E>. Ambient: <AUDCAP>Car horns, crowd murmur, distant train whistle<ENDAUDCAP>.选择アスペクト比:
16:9 或 9:16)。默认为 16:9。生成视频:
查看并迭代:
FAQ
Ovi 以 540p 分辨率生成视频,在画面质量与生成速度之间取得平衡。
Ovi 生成的每段视频片段时长约为 5 秒。
为了获得最佳结果,请加入详细的场景描述,使用 `<S>` 标签指定对白,并使用 `<AUDCAP>` 标签描述环境声音。清晰且有描述性的プロンプト有助于模型理解你的创意构想。
支持,你可以根据项目需求在 '16:9' 和 '9:16' アスペクト比之间进行选择。
Ovi 的定价具有竞争力,每次生成 $0.20;与同类模型相比,它能在不牺牲质量的情况下提供显著的成本优势。