← Back to all generators

lucataco/qwen2.5-omni-7b

Qwen2.5-Omni is an end-to-end multimodal model designed to perceive diverse modalities, including text, images, audio, and video, while simultaneously generating text and natural speech responses in a streaming manner.

Capabilities

Reference ImagesSystem Prompt

Cost

Community model (estimated from hardware time)

Input Parameters

audiostring

Optional audio input

generate_audioboolean

Whether to generate audio output

Default: true
imagestring

Optional image input

promptstring

Text prompt for the model

system_promptstring

System prompt for the model

Default: "You are Qwen, a virtual human developed by the Qwen Team, Alibaba Group, capable of perceiving auditory and visual inputs, as well as generating text and speech."
use_audio_in_videoboolean

Whether to use audio in video

Default: true
videostring

Optional video input

voice_typestring

Voice type for audio output

Default: "Chelsie"
ChelsieEthan
Version: 0ca8160f7aafUpdated: 7/25/202631.7K runs