September 8, 2026

What We've Been Making with LTX-2.5

Cartoon characters in my hiking photos, a second moon over the lake, and a lot of short experiments on our own hardware.

The fox walking down my hiking trail is one of my favorite things we’ve made with a local model. I like how consistent the character stays across different photographs, and especially how it walks along the path and interacts with the fire ring. We tried a bear too, taking the same idea into a few different scenes. I’m really impressed with how these turned out.

These started with photographs from my hikes. I wanted to see what would happen if we added something to a place I’d actually been, then asked LTX-2.5 to bring it to life. That led to camera experiments, creatures in the woods, and eventually a whole collection of prompts that had nothing to do with hiking.

I’ve put the results together below. The clips are grouped so you can watch several examples without an explanation interrupting every few seconds. The introductions use Ryan, a separate voice model running locally. The sound during the examples comes from LTX itself.

A cartoon fox from the LTX video experiments

Loading this video connects to YouTube and allows it to receive information about your visit.

Watch on YouTube
52 local LTX-2.5 experiments, with a Ryan introduction before each group. The examples retain their generated sound.

Download all 52 video prompts, image-edit prompts, and our setup guide. Use them as starting points for your own photographs and scenes.

Starting with my own photographs

With the fox and bear, we first used FLUX.2-klein to put the character into a photograph. Then we passed that edited still to LTX to generate movement and sound. Doing the image edit separately gave me a chance to decide whether the character looked right before we animated anything.

That is the part I keep coming back to: having a character I like and being able to use it in another scene. These are separate short generations, so I wouldn’t take them as proof that we can make a whole film with perfect continuity. But they’re good enough that I want to keep building on them. The interactions with the scene are what make the idea interesting to me.

We also tried moving through the photographs without adding a character. A little camera movement can do quite a lot with a still image. I picked a handful of these for the second group, followed by an early experiment that pans across a lake and reveals a huge moon. That one goes beyond camera movement, of course. It was one of the first generations, and it still belongs among my favorites.

The next group gets into adding subjects: a ferryman, Bigfoot, a stag, a wyrm, and a moose. I wanted the very first Bigfoot generation in here alongside some of the later attempts. Sometimes a prompt gives us something I want to keep straight away. Other times the subject or action doesn’t come out as intended, and another generation with the same basic idea doesn’t necessarily fix it.

For this kind of shot, editing the reference image first has been the most useful change to our process. It lets us deal with where the subject belongs and how it fits into the photograph before asking for movement. Some examples start from the original hiking photo, while others start from an edited picture. The moon, Bigfoot, and wyrm examples here use the original-photo approach; the ferryman, stag, and moose use edited references.

Seeing what else it could do

After the hiking experiments, I wanted to try a wider range of subjects. The rest of the video is a collection of short scenes generated from text: speech, different animation styles, action, physical effects, close-ups, and a few attempts at cuts within a single generation.

I like seeing these together because it gives me a much better feel for what I’d want to try next. A conversation, a moving car, and a product shot ask different things of the model. I don’t need to turn every one into a finished project to find out whether the result interests me.

We included prompts for hands, faces, reflections, and writing too. Those are places where I want to look closely before using a clip. A result can be enjoyable while still missing part of the request. The full prompts are in the download so you can compare what we asked for with what came back, including the attempts at two shots, three shots, and a match cut.

Most of these clips run for five seconds. The three cut experiments run for ten, and we’ve left all of them at their full generated length. That makes this a useful collection of small tests rather than a montage of only the best moment from each attempt. I’ve selected the examples I wanted to show, but within each example you get the whole take.

There is a group using the same scene prompt in different frame shapes as well. Those keep their original proportions in the compilation. The black space around a square or vertical example comes from fitting it into the video player, not from generating a different crop.

The setup behind the video

We’re running LTX-2.5 in ComfyUI on GX10 hardware with an NVIDIA GB10 and 128GB of unified memory. We have two machines available, and a shared dispatcher sends jobs to whichever is free. Each of these generations ran on one machine. We aren’t splitting one video across the pair.

The transformer is the distilled 22-billion-parameter model in NVFP4. All the clips in this compilation used the int8 text encoder. That matters because our earlier experiments also included a larger bf16 encoder, and mixing their settings together would make the instructions confusing. The settings below come from these actual video files.

What I enjoy about having this running locally is being able to keep trying things. The hardware has a cost, but once it’s set up I can queue another idea without deciding whether that particular attempt is worth paying a generation fee. Short clips make that process manageable. Across 460 five-second image-to-video runs, the median was 3:11, with a range of 2:46–3:47; the middle half took 3:07–3:18.

The workflow generates video and audio, enlarges the video latent with a separate upscaler, runs a second sampling stage, and decodes the result. The guide collects our settings and the steps we follow when trying a new scene.

SettingUsed in this showcase
Video modelLTX-2.5 distilled 22B, NVFP4
Text encoderGemma 4 LTX int8 convrot
Usual base resolution768×576
Usual generated output1536×1152 after 2× latent upscale
Frames121 for most clips; 241 for the three cut experiments
Frame rate24 fps
Seed42
Video / audio guidance1 / 1
SamplerEuler ancestral
Sampling stages8 steps, then 3
Video decodingTiled

The four framing experiments use different dimensions, recorded individually in the download. The compilation itself is delivered at 1920×1080 with the original proportions preserved. That delivery size isn’t the resolution at which the examples were generated.

How long the generations took

These are historical measurements from our retained logs, timed from prompt submission through completion. They cover a mix of experiments and encoder settings, rather than a controlled comparison or a timing guarantee for the selected clips. The five- and ten-second labels mean 121 and 241 frames at 24 fps.

Historical LTX generation times: boxes show the middle half, lines the full range, and marks the median. Exact timings and counts follow in the table.

RunnMedianRangeMiddle halfBase → outputPeak unified memory used, GiB (median; range)
Image → video, 5s4603:112:46–3:473:07–3:18768×576 → 1536×115285.8; 84.3–105.0
Image → video, 10s146:376:26–6:576:30–6:51768×576 → 1536×115291.8; 87.8–91.9
Text → video, 5s282:582:54–3:102:57–3:00768×576 → 1536×115287.2; 85.2–89.1
Text → video, 10s36:106:04–6:126:07–6:11768×576 → 1536×115291.8; 88.0–91.8

Times are minutes:seconds, rounded to the nearest second. The table groups the usual resolution; other frame shapes and longer experiments are excluded. We kept successful runs and excluded likely cached results whose peak memory rose less than 1 GiB above the starting level. The ten-second text-to-video group has only three runs.

Peak unified memory used is the highest sampled whole-system memory use during each run. GX10 memory is shared by the CPU and GPU, so these readings include the model, generation workload, and other system processes. The memory column gives the median and range of those per-run peaks. These logs do not contain GPU utilization, power, or bandwidth measurements.

The rotating-watch example is a concrete reference from the text-to-video group: 177.5 seconds and 85.5 GiB peak unified memory used. It used seed 42, 121 frames at 24 fps, and a 768×576 base enlarged to 1536×1152 output, with the showcase settings above.

Try it with a photograph of your own

The short setup and process guide accompanies the video and image-edit prompts. Pick a photograph with space for something to happen, then decide whether you want to move through the existing scene or add a subject. For an added character, use an image-edit prompt as a starting point and check the still before animating it.

Describe what you want to happen, how the camera should move, and what should be audible. If somebody speaks, write their lines into the prompt. Then watch the result against that request before deciding what to change. When the problem is the placement or appearance of an added character, go back to the still image and fix it there.

I’m especially interested in doing more with the recurring characters. Having the fox show up in photographs from places I’ve been is already enough to keep me experimenting. Open a photograph from somewhere you know, give it a small action, and see what comes back. If you add a character you like, try taking it into a second photograph.

What We've Been Making with LTX-2.5
0:00
0:00