← Back to all generators
lucataco/qwen2.5-omni-7b
OfficialView on Replicate →
Qwen2.5-Omni is an end-to-end multimodal model designed to perceive diverse modalities, including text, images, audio, and video, while simultaneously generating text and natural speech responses in a streaming manner.
Capabilities
Reference ImagesSystem Prompt
Cost
Community model (estimated from hardware time)
Input Parameters
| Name | Type | Description | Default | Constraints |
|---|---|---|---|---|
audio | string(uri) | Optional audio input | — | — |
generate_audio | boolean | Whether to generate audio output | true | — |
image | string(uri) | Optional image input | — | — |
prompt | string | Text prompt for the model | — | — |
system_prompt | string | System prompt for the model | "You are Qwen, a virtual human developed by the Qwen Team, Alibaba Group, capable of perceiving auditory and visual inputs, as well as generating text and speech." | — |
use_audio_in_video | boolean | Whether to use audio in video | true | — |
video | string(uri) | Optional video input | — | — |
voice_type | string | Voice type for audio output | "Chelsie" | ChelsieEthan |
audiostringOptional audio input
generate_audiobooleanWhether to generate audio output
Default:
trueimagestringOptional image input
promptstringText prompt for the model
system_promptstringSystem prompt for the model
Default:
"You are Qwen, a virtual human developed by the Qwen Team, Alibaba Group, capable of perceiving auditory and visual inputs, as well as generating text and speech."use_audio_in_videobooleanWhether to use audio in video
Default:
truevideostringOptional video input
voice_typestringVoice type for audio output
Default:
"Chelsie"ChelsieEthan
Version:
0ca8160f7aafUpdated: 7/25/202631.7K runs
cinemasetfree.com