Guides

LTX-2 Prompts: A Working Guide to LTX-2.5

LTX-2 prompts are prose, not tags. The six elements LTX documents, worked text-to-video, image-to-video and LTX-2.5 multishot examples, plus ComfyUI settings.

By Aditya Bawankule13 min read

LTX prompts are written as prose, not as tags. LTX's own prompting documentation asks for six things in one flowing paragraph: the shot, the scene, the action, the character, the camera movement, and the audio, in present tense, with any spoken line in quotation marks. There is no token syntax, no weight brackets, and no quality-tag tail. Getting an LTX-2 prompt right is mostly a matter of writing a paragraph a cinematographer could shoot.

LTX-2.5 shipped on August 11, 2026 and adds one genuinely new prompt-craft job: native multishot, where a single prompt describes several connected shots joined by cuts. That changes what a prompt has to carry. This guide covers the format LTX documents, worked text-to-video, image-to-video, and multishot prompts with the reasoning behind each, the failure modes that waste a generation, and the ComfyUI parameters that matter if you run the weights yourself. Every prompt below also runs as written in Dream Pixel Forge, which ships LTX-2.5 in the browser, so you can test the craft without a 32GB card.

Which LTX are you prompting: LTX-2, LTX-2.3, or LTX-2.5

Three versions are in circulation and the naming matters more for access than for prompt grammar.

  • LTX-2 is the open-weights audio-video model Lightricks released in January 2026. On the hosted API it was deprecated on July 15, 2026: requests naming ltx-2-fast or ltx-2-pro are already served by LTX-2.3 at LTX-2 prices, and the model ids are removed on August 15, 2026.
  • LTX-2.3 is the version most third-party tools and tutorials are still built on, and it remains the API fallback for anything that names LTX-2.
  • LTX-2.5 is the current release: a 22B distilled transformer with a custom Gemma 4 12B text encoder, a diffusion video decoder, automatic duration prediction, native 4K HDR and EXR output, and native multishot.

The prompt grammar is shared. The six elements below come from LTX's official prompting guide, which covers LTX models generally, so a good LTX-2.3 prompt is a good LTX-2.5 prompt. Only multishot is version-specific: LTX-2.5 generates several shots inside one clip, and earlier versions do not.

What is the best way to prompt for LTX-2?

Write one paragraph that covers six elements, in roughly this order. LTX documents them as the things the model needs to build the scene.

  1. Establish the shot. Use real cinematography terms, and include shot scale. “A low-angle medium shot” is instruction; “a beautiful shot” is not.
  2. Set the scene. Lighting condition, color palette, surface texture, atmosphere. LTX's guidance is explicit that one coherent light logic per shot works and mixed light sources confuse the result.
  3. Describe the action as a natural sequence that flows from beginning to end, in present tense.
  4. Define the characters: age, hairstyle, clothing, distinguishing features. Express emotion through physical cues, not labels. Write “his jaw tightens and he looks away” rather than “he is sad.”
  5. Identify the camera movement, including when it happens. The guide adds a detail worth internalizing: describing how subjects look after the move helps the model complete the motion instead of abandoning it halfway.
  6. Describe the audio. Ambient sound, music, speech, or singing. Spoken dialogue goes in quotation marks, and you can specify the language and accent.

Length is matched to complexity rather than capped. For a single continuous take LTX suggests roughly four to eight descriptive sentences as a single flowing paragraph. That is a change from the older LTXV 0.9.x guidance in the LTX-Video repository, which told you to keep prompts within 200 words; the current advice is that a longer prompt is fine provided every sentence adds concrete visual or audio detail. Sentences that add adjectives without adding information are the ones that hurt.

The other structural choice is prose versus screenplay. A single continuous take is a paragraph. A scene with dialogue, multiple beats, or precise timing can be written as a screenplay, with scene headers, character cues, and quoted lines, which is the format LTX uses for its own sample prompts. The fundamentals do not change between the two: present tense, physical emotion cues, dialogue in quotes.

LTX video prompts you can copy

Three worked examples, one per input mode, with the reasoning that makes each one work.

Each one was then run through LTX-2.5 exactly as printed, and the clip it produced sits directly under it, so you can check the craft advice against the result rather than taking our word for it. Nothing was re-rolled for a better take: every clip on this page is the first render of the prompt above it. Turn sound on — LTX generates the audio in the same pass as the picture, and a muted clip hides half of what these prompts control.

Text-to-video: a single continuous take

A slow push-in on a woman in her forties, gray-streaked hair tied back, wearing a worn canvas apron, standing at a steel workbench in a narrow ceramics studio. Late-afternoon light comes through a single dusty window on the left, warm on her hands and falling off into deep shadow behind her. She lifts a half-finished bowl, turns it once against the light, and her shoulders drop as she sets it back down. The camera settles on the bowl as she steps out of frame to the right. Ambient sound: the low hum of a kiln, a distant radio, no music.
LTX-2.5, five seconds, 16:9, 720p, generated from the prompt above with no reference image. The whole beat survives: she lifts the bowl, turns it once against the light, sets it back down, and the camera stays with the bowl as she leaves frame to the right. That last move is the part a prompt most often loses, and it is the one the closing sentence spends its words on. The single window on the left gives the shot real shadow falloff rather than an even studio wash. One instruction it ignored: “steel workbench” came back as stone.

Why it works: one shot scale and one camera move, named once. One light source with a stated direction, which satisfies the coherent-light-logic rule. One beat, with the emotion expressed physically as the shoulders dropping rather than labeled as disappointment. The final sentence tells the camera where to end up after the move, and the audio is named including the absence of music, so the model does not invent a score.

Image-to-video: describe what happens, not what it looks like

The man turns his head toward the camera and breaks into a slow grin, steam still rising from the mug in his hands. The camera pushes in gently to a chest-up framing and holds. Ambient sound: rain against the window, the clink of the mug on the table, no music.
The still the prompt above animates. It has to ship with the clip, because an image-to-video prompt is only judgeable against the frame it was given: everything the prompt refuses to mention — the room, the sweater, the rain on the glass, the cool overcast grade — was already decided here.
Five seconds, 16:9, animating the still above. The first sentence lands exactly: he turns, the grin arrives slowly, the steam keeps rising. The room, the sweater, the rain on the glass, and the cool grade all hold, which is the argument for not re-describing them. The second sentence only half lands. The camera does push in, but he stands up while it does and his head leaves the top of the frame, so the shot never settles on the chest-up framing it was told to reach. That is this page's own camera rule failing from the other direction: the prompt said where the camera should end up and never said where he should be when it got there.

Why it works: it says nothing about the room, the wardrobe, or the grade, because the source image already decided all of it. LTX's image-to-video guidance is that the prompt describes what should happen, focusing on motion and action, camera movement, and audio, since the model already knows what the scene looks like. Re-describing the frame in an image-to-video prompt is the most common way to make the output drift away from the image you approved.

Multishot: one prompt, several cuts

This is LTX-2.5's own documented example, and it is worth reading as a template rather than paraphrasing it:

A wide shot frames a rainy city intersection at dusk, neon signs reflecting on wet asphalt. A young woman in a yellow raincoat walks toward camera, gripping a folded newspaper, while cars hiss past behind her. Soft synth music and distant traffic fill the air. A hard cut transitions to a medium close-up of her face under the hood, raindrops catching the neon as she looks off-screen left; the synth score continues across the cut, traffic muffled. She whispers, “He's late.” Another hard cut jumps to a low-angle shot of a man's scuffed boots stepping into a puddle at the curb; the music drops to a low drone. He lifts his head into frame, short dark hair, soaked jacket, and smiles toward her off-screen as a bus rumbles past.
Ten seconds, 16:9, first render of LTX's own example prompt. Ten rather than five, set by hand, because three shots do not fit into five seconds — which is the practical case for letting the model choose the duration on a multishot prompt instead. Both cuts land, at 2.96s and 6.24s, and each new shot is the one the prompt named: the wide intersection, then her face under the hood, then the boots at the curb before he lifts his head into frame. She is recognizably the same woman across the first cut, which is what re-identifying her on the far side of it buys. Transcribed, the whispered line comes back as “He's late.” And the neon is unreadable gibberish, which is the on-screen-text limit listed further down this page showing up in LTX's own showcase scene.

Three shots, and every cut does four things: it names the transition in plain language (“a hard cut transitions to”), re-establishes shot scale and framing, keeps the subject identifiable by reusing her visual identifiers, and states what happens to the audio across the cut. Those four moves are the whole multishot technique.

LTX ships this frame under its prompt-adherence claim for LTX-2.5: shorter prompts, more of the instruction retained. Source: the LTX-2.5 model page.

How to write LTX-2.5 prompts for multishot without losing continuity

Multishot has its own rules, and they are the opposite of what most prompt templates teach. LTX asks for the whole scene as one chronological paragraph and explicitly says not to use a shot list, numbered beats, or screenplay sluglines unless you also describe the cut in prose. A numbered list of shots reads to the model as a list, not as an edit.

ElementSingle shotMultishot
CameraOne continuous takeNew framing after every cut
TransitionsCamera moves onlyName the edit: hard cut, match cut, dissolve
ContinuitySame space and subjects throughoutRe-identify subjects when they reappear
AudioOne continuous soundscapeSay at each cut whether sound continues or changes

Two practical limits come straight from the documentation. Prefer two to four shots in one generation, because more cuts need shorter, clearer beats per shot than a single prompt can usually carry. And avoid conflicting geography or unexplained costume changes between cuts unless you say the cut jumps time or place. Chronology markers help: “initially,” “a moment later,” “simultaneously.”

Stay single-shot when you want unbroken camera motion, an intimate performance, or dialogue that has to stay lip-synced in one framing. And when you are starting from a first frame, stay single-shot unless you deliberately want to cut away from the image you supplied.

Multishot is also the fastest way to find out whether your prompt writing is the problem or your setup is, and that is easier to test on a hosted surface than in a node graph. Dream Pixel Forge runs LTX-2.5 for text-to-video and image-to-video at 720p and 1080p, 5 to 20 second clips with synchronized audio, on the same credit balance as its other video engines, so you can paste the examples above and iterate on wording before deciding whether local inference is worth the hardware.

How to prompt LTX i2v, including first and last frame

Image-to-video is where prompt discipline pays most. The ComfyUI template injects your source image into the video latent at strength 0.7 in the first stage, then re-injects it at strength 1.0 at full resolution, so the composition is anchored twice. Everything you spend words on that is already visible in the frame competes with that anchoring.

A useful shape for an image-to-video prompt is three clauses: what the subject does, what the camera does, what you hear. LTX's own example is one sentence long: the woman turns to face the camera and smiles, a warm breeze moving through her hair, soft piano music in the background.

The first-frame / last-frame template (video_ltx2_5_flf2v) generates the motion between two images you supply. It is single-stage, so there is no 2x upscale and the output stays at the base resolution you set, and the prompt should describe the journey between the frames rather than either endpoint. On the hosted API the equivalent is last_frame_uri, which cannot be combined with automatic duration, since a fixed ending needs a known length.

The general argument for animating an approved still rather than rolling a scene from text is the same on every engine, and it is laid out in our guide to image to video AI.

Prompting dialogue, audio, and dubbing

LTX generates audio jointly with video: in the ComfyUI pipeline an empty audio latent is concatenated with the video latent and the two are sampled together, which is what keeps them in sync. Because sound is generated in the same pass, it is promptable, and leaving it unspecified means the model decides for you.

Three habits cover most of it. Put spoken lines in quotation marks and name the language or accent when it is not obvious. Name ambience explicitly, including what you do not want, because an unrequested score is the most common audio surprise. And keep a spoken line short enough to fit the clip at a natural pace.

All three habits in one prompt:

A handheld medium close-up of a woman in her late twenties, dark hair pushed behind one ear, standing in a tiled corridor lit only by the fluorescent tube directly above her. She looks straight into the lens, hesitates, then lifts her chin as she speaks. The camera holds a chest-up framing and drifts in slightly. She says, in an English accent, “I found it. It was under the stairs.” Ambient sound: a fluorescent hum and a door closing far down the corridor, no music.
Five seconds, 9:16, 720p, first render. Run the audio through a transcriber and it comes back as “I found it. It was under the stairs.” — the quoted line word for word, at a pace that fits the clip because the line was written to fit it. Underneath it is the hum and the corridor the prompt asked for and no score, which is what naming the absence of music is for. Two things it took liberties with: the drift-in is a good deal stronger than “slightly,” and her face shifts a little as the camera closes on it.

Speech replacement in existing footage is a separate tool with its own format. The Dub-It IC-LoRA takes a source video and a prompt in a fixed template: [Speaker] is speaking [Language/Accent], saying: “[Dialogue]”. Three constraints are easy to trip over: it does not translate for you, so supply the full target-language text; write it in the native script, Cyrillic for Russian and Chinese characters for Mandarin; and the beta adapter handles one speaker only. Match the syllable length of the original, since a prompt that is too long makes the model skip words and one that is too short produces unnaturally slow delivery. LTX lists English, French, Spanish, German, and Russian as validated.

Parameter guidance for running LTX-2.5 in ComfyUI

Prompt craft only matters once the generation actually runs. LTX ships native ComfyUI templates: open the Templates browser, search LTX-2.5, and pick video_ltx2_5_t2v or video_ltx2_5_i2v, then use Download all to fetch the distilled 22B transformer, the Gemma 4 text encoder and prompt enhancer, and the video and audio VAEs.

The settings that change your results:

  • Width and height must be divisible by 32. The template defaults to 768 by 512, and the pipeline upscales 2x in the second stage, so the file you get is 1536 by 1024. Set the base resolution to half of your target, not the target.
  • Length must be one plus a multiple of 8 frames. The default 97 frames at 24 fps is about four seconds. 24 fps is cinematic, 25 is broadcast standard, 30 is smoother.
  • Prompt Enhance is on by default. A dedicated Gemma 4 enhancer expands a short prompt into a longer one before encoding. That is useful when you are exploring and actively unhelpful when you are testing whether a specific phrasing works, because you are no longer measuring your prompt. Turn it off when you iterate on wording.
  • A negative prompt is already applied by the template: pc game, console game, video game, cartoon, childish, ugly. If you actually want a stylized or animated look, that default is working against you, so edit it before blaming your prompt. The same reasoning applies across models, and our guide to the negative prompt covers when one earns its place.
  • Distilled versus full. The templates load the distilled model, which is built to produce good results in few steps and is the right place to iterate on prompts. The full model, used in the two-stage generation workflow, takes more steps for finer detail and more nuanced motion. Find the prompt on distilled, spend the time on full.
  • Seed. Stage 1 randomizes by default. Fix it when you want to compare two prompt phrasings, otherwise you are comparing two different rolls.

Hardware is the real gate on self-hosting. LTX documents a minimum of an NVIDIA GPU with 32GB or more of VRAM, 32GB of system RAM, 100GB of free storage, CUDA 12.7 or higher, and Python 3.12 or higher, and it recommends an A100 80GB or an H100. That is data-center hardware, not a typical desktop card, so verify the requirement against your machine before planning a local workflow around it.

Paying per second instead is the other route. LTX's own API runs ltx-2-5-fast up to 4K and ltx-2-5-pro up to 1080p, in 16:9 and 9:16, with published rates starting at 0.09 dollars per second at 720p on Fast and 0.12 on Pro. Clips longer than 10 seconds are available at 720p and 1080p at 24 or 25 fps; the higher resolutions and the 48 and 50 fps modes cap at 10. Sending "duration": null lets the model pick the clip length from your prompt, which pairs well with multishot, where the right length depends on how many cuts you described.

Common LTX prompt failure modes

  • Tag-stack prompts. LTX takes prose. A comma-separated pile of style words carries far less information than one sentence of shot description, and there is no weighting syntax for it to parse.
  • Mixed lighting. Two competing light sources in one shot description is the failure LTX calls out by name. Pick one light logic per shot.
  • Numbered shot lists for multishot. A list is not an edit. Describe the cut in prose or the model reads your beats as a single confused scene.
  • Naming a subject only once in a multishot prompt. If a person reappears after a cut and you do not re-identify them, identity is unlikely to survive the cut.
  • Re-describing the scene in an image-to-video prompt. The frame is already decided at conditioning strength 0.7 and then 1.0. Prompt motion, camera, and sound only.
  • Relying on on-screen text. LTX-2.5 improves short-text accuracy, but its own documentation says exact spelling and frame-to-frame consistency are not guaranteed. Keep text short and prominent, check it across the whole clip, and add critical titles and logos in post.
  • Chaotic physics. Highly chaotic motion still introduces artifacts. Simpler, plausible motion is more reliable, which is a prompt decision as much as a model limit.

The first of those is worth seeing rather than taking on faith. Here is the ceramics scene from the top of this page rewritten as a tag stack, keeping every noun and adding the usual quality-tag tail, run at the same length, aspect ratio, and resolution as the prose version:

woman, 40s, gray-streaked hair, canvas apron, ceramics studio, steel workbench, half-finished bowl, window light, push-in, cinematic, 8k, masterpiece, ultra detailed, dramatic lighting, bokeh, film grain, award winning photography, best quality, trending
Five seconds, 16:9, 720p. This is the honest version of the failure, and it is not a broken clip. The tag stack does not confuse LTX: every noun is on screen and the frame is pretty. What it does not buy is a scene. Nothing happens — she holds the bowl for five seconds and the camera drifts — because no tag describes an action with a beginning and an end. The light is soft and sourceless where the prose version has one window and a hard falloff, because “dramatic lighting” names a mood and “through a single dusty window on the left” names a direction. “Steel workbench” is ignored here too. Play it against the first clip on this page: the difference is not quality, it is information.

Running the same prompts without a GPU

Nothing above except the ComfyUI parameters depends on where the model runs. Dream Pixel Forge ships LTX-2.5 alongside Google Veo 3.1, ByteDance Seedance, and xAI Grok Imagine Video on one credit balance, so the prompts in this guide go straight into the browser: text-to-video and image-to-video, 720p and 1080p, 5 to 20 second clips with synchronized audio, no ComfyUI install and no CUDA version to match.

Having the engines side by side is also the fastest way to feel how much the dialects differ, which matters the moment you paste a prompt across. Veo 3 prompts use plain scene description with no token syntax and cap at 8 seconds. Seedance prompts use numbered reference tags to assign roles to up to nine reference images, which LTX has no equivalent for. Grok animates a single photoreal still where the others refuse. LTX is the one that will cut between shots inside a single generation. For the wider field, our roundup of AI video models in 2026 compares them on access and output, and open source AI video generators covers what you can actually self-host.

Try it now

Free to try, no account needed

Reference images

Select up to 8 images to guide the result.

0/8
Example output for Freeform Image Generator
Example output

Your generated image will replace this example.

The short version

Write one paragraph in present tense covering shot, scene, action, character, camera, and audio. Keep one light logic and one clear beat per shot. Put dialogue in quotes. For multishot, name each cut in prose, re-establish the framing, re-identify the people, and say what the sound does across the cut. For image-to-video, prompt only what changes. Iterate on the distilled model with the prompt enhancer off, then commit to the full model once the wording is right.

If you want a deeper look at what LTX-2.5 changed under the hood, our write-up on LTX-2.5 covers the release itself. And if the shot you want starts from a still, generating that frame first is still the cheapest way to get a good clip: the freeform generator gets you one in seconds.

Tools for this guide

Frequently asked questions

What is the best way to prompt for LTX-2?

Write one flowing paragraph in present tense covering six elements: the shot (using real cinematography terms and a shot scale), the scene (lighting, color palette, texture, atmosphere), the action as a natural sequence, the characters (age, hair, clothing, distinguishing features, with emotion expressed through physical cues rather than labels), the camera movement including when it happens, and the audio, with any spoken line in quotation marks. LTX takes prose, not tags: there is no token syntax, no weighting brackets, and no quality-tag tail to add. Keep one coherent light logic per shot, since mixed light sources are the failure LTX calls out by name.

How to write LTX-2.3 prompts?

The same way you write LTX-2.5 prompts. LTX's prompting guidance covers the model family rather than a single version, so the six-element paragraph, present tense, physical emotion cues, and quoted dialogue all carry across LTX-2, LTX-2.3, and LTX-2.5. The one difference is multishot: LTX-2.5 can generate several connected shots joined by explicit cuts inside a single prompt, and earlier versions cannot, so a multishot prompt sent to LTX-2.3 is read as one confused scene. On the hosted API, LTX-2 was deprecated on July 15, 2026 and is removed on August 15, 2026, with requests served by LTX-2.3 in the meantime.

How to prompt LTX-2.3 i2v?

Describe what should happen, not what the scene looks like. In image-to-video the source frame has already decided composition, wardrobe, lighting, and grade, so the prompt should carry only motion and action, camera movement, and audio. A useful shape is three clauses: what the subject does, what the camera does, what you hear. Re-describing the scene fights the conditioning and is the most common cause of drift away from the frame you approved. If you supply both a first and a last frame, describe the journey between them rather than either endpoint.

What are some good prompts?

The shots worth having templates for are a single continuous take (one subject, one camera move, one light source, one beat, audio named including the absence of music), an image-to-video motion-only prompt, and an LTX-2.5 multishot scene of two to four cuts. This guide has a complete worked example of each, plus LTX's own documented multishot sample. For multishot, every cut needs four things: the transition named in plain language, the new framing re-established, recurring people re-identified by their visual details, and a statement of what the audio does across the cut.

Can I use LTX-2.5 prompts without a GPU?

Yes. Dream Pixel Forge runs LTX-2.5 in the browser for text-to-video and image-to-video at 720p and 1080p, in 5 to 20 second clips with synchronized audio, on the same credit balance as Veo 3.1, Seedance, and Grok Imagine Video, so the prompts in this guide run as written with no ComfyUI install and no CUDA version to match. LTX also sells hosted API access per second of output, and the weights themselves are free to use commercially for organizations under 10 million dollars in annual recurring revenue.

What hardware do I need to run LTX-2.5 locally?

LTX documents a minimum of an NVIDIA GPU with 32GB or more of VRAM, 32GB of system RAM, 100GB of free storage, CUDA 12.7 or higher, and Python 3.12 or higher, and recommends an A100 80GB or an H100. That is data-center hardware rather than a typical desktop card. The ComfyUI templates soften the requirement by generating at a low base resolution and upscaling 2x in a second stage, with a default of 768 by 512 at 97 frames, but the documented floor is still the number to check your machine against before planning a local workflow.