Use Text-to-Video to create a new video from words, or AI Image Generator for a still image. If an existing image should guide the result, start with Image-to-Video or AI Image Editor.
Create from text
Open the tool and sign in.
Write what should appear: the subject, action, setting, and visual style.
Choose the model and the available duration, resolution, aspect ratio, and output quantity. Options differ by model and plan.
For video, review the Audio control. Some models generate sound; uploaded audio is a separate model-specific option.
Review the displayed credit estimate, then generate.
Open the completed result in History or My Library, check it, and download it.
A useful prompt structure
Subject + action + setting + camera or composition + lighting or style. Start with one clear idea. Add details that matter, and remove instructions that contradict each other. Use short, concrete sentences instead of a long list of loosely related keywords.
Goal | Example prompt |
Product image | A matte blue ceramic mug on a pale wooden table. Soft window light from the left. Clean background, close-up product photography. |
Video with simple motion | A cyclist rides along a quiet coastal road at sunrise. The camera follows steadily from behind. Gentle movement, natural light. |
Animate a reference image | The person smiles and turns slightly toward the window. Keep the camera still and the background unchanged. |
These examples guide the model; they do not guarantee exact text, facial identity, physical motion, or every requested detail. Use the aspect-ratio control for the output shape instead of relying only on words in the prompt.
Add sound or your own recording
Some Text-to-Video models offer generated audio through an Audio switch. For your own MP3 or recorded voice, select LTX-2.3 and use Drop audio or browse. Its Add Audio picker also offers Text to Speech, voice cloning, and presets. Follow the audio-upload guide for supported inputs and reference restrictions. It is no longer accurate to say Text-to-Video cannot accept uploaded audio.
Make a longer story, poem, or music video
Break the idea into short scenes. Give each clip one main action.
Choose a duration supported by the selected model; a plan allowance expressed in hours is not the maximum length of a single clip.
Use consistent reference material or a saved Character where supported.
Generate a short test before spending credits on the full sequence.
Review each clip, then assemble clips and any exact captions or soundtrack in an editor that supports the timing you need.
For extending an existing clip, try AI Video Extender. For a portrait speaking a longer script, see Talking Photo and its mode-specific limits.
Improve a result that misses the prompt
Keep the main subject and requested action near the beginning.
Specify which parts should change and which should stay the same.
Avoid asking for several unrelated camera moves or scene changes in one short clip.
Check that the selected model supports your inputs. Reference attachments can change how the starting image is created.
Compare the same input and settings when assessing a change. Each retry uses the displayed credits.
If Generate is disabled or an output is broken
Read the inline validation, check required fields and completed uploads, and confirm the current estimate is within your balance. Uploaded audio cannot currently be combined with reference attachments in LTX-2.3. If the same supported request repeatedly fails, stop retrying and send support the route, model, project link if available, error, and screenshot. See disabled Generate buttons and failed or technically broken results.
