# LTX showcase: prompts and process

`prompts.tsv` contains all 52 video prompts in showcase order. `image-edit-prompts.tsv`
contains the edit prompts for the nine edited references, labeled by example. Adapt
those descriptions to your own photographs, subjects and surroundings.

## Start with your own photograph

1. Choose a photograph with room for the action you want to create.
2. If adding a character or object, edit the reference image first. We used FLUX.2-klein
   4B NVFP4 for this stage. The image-edit prompts show how we described the additions.
3. Check the subject's appearance, scale and placement in the still. Fix those before animation.
4. Use the still as the image-to-video reference in ComfyUI. Describe the movement,
   camera direction and sound you want. Write any spoken lines into the video prompt.
5. Watch the whole result and compare it with the request. Change the still for placement
   problems; change the video prompt for movement or sound, then try another take.

For camera-only movement, we started directly from a photograph. The moon, Bigfoot and
wyrm examples also started from original photos; the cartoon characters, ferryman, stag
and moose started from edited references. Text-to-video examples, including the watch,
start from the written scene description.

## Our setup

Each generation ran on one GX10 with an NVIDIA GB10 and 128GB of unified memory. A shared
dispatcher assigned jobs across two machines. We used ComfyUI with LTX-2.5 distilled 22B
NVFP4 and the Gemma 4 LTX int8 convrot text encoder for the selected clips, with the
LTX-2.5 video/audio VAEs and 2× latent spatial upscaler. The workflow generates video and
audio, upscales the video latent, samples again, then decodes.

| Setting | Showcase value |
|---|---|
| Usual base → output | 768×576 → 1536×1152 |
| Frames / frame rate | 121 at 24 fps; 241 for the three `BR-ms-*` cut examples |
| Seed | 42 in both stages |
| Video / audio guidance | 1 / 1 |
| Sampler | Euler ancestral |
| Sampling stages | 8 steps, then 3 |
| Video decoding | Tiled |

The framing experiments use these dimensions:

| Example | Base → output |
|---|---|
| BR-ar-16x9 | 896×512 → 1792×1024 |
| BR-ar-cine | 960×416 → 1920×832 |
| BR-ar-1x1 | 768×768 → 1536×1536 |
| BR-ar-9x16 | 512×896 → 1024×1792 |

The compilation fits each source into 1920×1080 with its proportions preserved. That is
its delivery size. The table gives the generated dimensions.

These are starting points for experiments with your own pictures. Historical software
and model revisions were not fully retained, and another installation can produce different
results. Try a short clip first and check the generated action and speech before using it.
