Yes—you can use your own audio in Text-to-Video and Image-to-Video. The current uploaded-audio workflow is available with LTX-2.3. An Audio toggle on another model can enable generated sound without enabling file uploads.
Upload a recording
Open Text-to-Video, or Image-to-Video if you have a starting image.
Open the model selector and choose LTX-2.3.
Keep Audio enabled, then select Drop audio or browse.
In Add Audio, choose Upload from device and select your file. MP3 is a supported audio format.
Wait for the upload and choose the audio section. Review the selected duration and any validation message.
Describe the scene or motion if needed, choose the supported resolution, and check the credit estimate before generating.
Why is there no upload button on LTX-2.5?
LTX-2.5 currently shows an Audio switch for generated audio. It does not expose the LTX-2.3 upload control. Select LTX-2.3 when your own recording is required, or use Audio-to-Video. Model availability and controls can change; check the options shown in your editor.
Generate or clone the voice inside the audio picker
The LTX-2.3 Add Audio picker also offers Text to Speech, Clone your voice with AI, and presets. If you already have the correct recording, upload it directly. If you need a reusable voiceover first, create and download it with AI Voice Generator or AI Voice Cloner.
References, start images, and end images
Uploaded audio cannot currently be combined with character references or additional reference attachments in these LTX-2.3 workflows. Remove the conflicting references to continue.
In Image-to-Video, you can use the starting image with audio, but an end image is not available with uploaded audio.
For multiple pictures, exact scene order, or a full song, plan separate clips and assemble them afterward. A paid plan does not make an unsupported input combination available.
If a control becomes unavailable, read its inline explanation before generating. Changing a model can change supported inputs, duration, resolution, and price.
Choose the workflow that matches your goal
You want | Use |
A still face to speak or sing | |
A face in an existing video to follow new audio | |
A new scene or animated starting image driven by audio | LTX-2.3 in Text-to-Video or Image-to-Video |
Visuals from audio without an existing video | |
Music-reactive animation effects |
Generated motion and dialogue can vary. Uploading a song does not guarantee exact choreography, frame-perfect timing, or commercial rights to the music. For a failed upload or broken result, send support the route, model, selected audio length, visible error, and project link if available.
