Instructions to use LiconStudio/LTX-2.3-Multiple-Subject-Reference with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Diffusers
How to use LiconStudio/LTX-2.3-Multiple-Subject-Reference with Diffusers:
pip install -U diffusers transformers accelerate
import torch from diffusers import DiffusionPipeline # switch to "mps" for apple devices pipe = DiffusionPipeline.from_pretrained("LiconStudio/LTX-2.3-Multiple-Subject-Reference", dtype=torch.bfloat16, device_map="cuda") prompt = "Astronaut in a jungle, cold color palette, muted colors, detailed, 8k" image = pipe(prompt).images[0] - Notebooks
- Google Colab
- Kaggle
Thank you. I am using this to cerate series.
I am combining msr v2, with image to video and image to image video, i use msr as the foundation. I would like to ask about memory, right now i am using a amd 5950x23d with 64gb, and a 5070 ti with 16gbvram, and i did some tweaking to use 16bit diffusion, any how, i cannot get more than 2 characters to work at all. is this a limitation of ltx, my 16gb, or msr? what i mean as well is the characters don't render and become clones or something else, clothing isn't right, and dialogue fails... i have managed to get a series going with just 2 people and editing.. I would like to know if i got 32gb vram would i be able to use 3 characters or more?
I am combining msr v2, with image to video and image to image video, i use msr as the foundation. I would like to ask about memory, right now i am using a amd 5950x23d with 64gb, and a 5070 ti with 16gbvram, and i did some tweaking to use 16bit diffusion, any how, i cannot get more than 2 characters to work at all. is this a limitation of ltx, my 16gb, or msr? what i mean as well is the characters don't render and become clones or something else, clothing isn't right, and dialogue fails... i have managed to get a series going with just 2 people and editing.. I would like to know if i got 32gb vram would i be able to use 3 characters or more?
It’s a bit of both. With more than two characters, things can occasionally get mixed up, but I rarely see this on my PRO 6000. Layer-by-layer loading does cause some loss in quality. I’d also suggest trying my LTX2.5-MSR, which has fewer issues with unwanted object duplication.
I am using msr v2... and yes it is a bit better than v1... which i was using before. I still only get two reliable characters with good descriptions in the top, if do 3 it falls apart even the dialogue.... So with 96 gb of ram you rarely see issues... I have a some what stable 16 bit difusion, 16 bit gema cpu , and i get about stable 20 second in 1280x672 with rtx upscaler at the end up to 2k. I also turned off the upscaler on the msr, and just use the first pass with a vae modification. i am able to spit out enough content its going good. 20 second clips at 16 bit... man the rtx pro 5000 - 72gb is what i wish i could get but its 10k and i cant, was thinking about the rtx pro 4500 but that's 5k... kinda sucks. but i would love to be able to have 5 people in the msr work with dialogue and action sequences...
Also i was going to code my own msr style thing to copy what h3 does, because the h3 system is much better, its built in, and i stopped using h3 because i am an American and i cannot release my stuff on youtube with them. and ltx is much better visually to my eyes.
I previously experimented with something I called VAC Decypher. The idea was to track and lock the tokens or conditioning associated with each character, so their identities, clothing, hair, and physical features wouldn't bleed into other characters.
I was trying to achieve something closer to H3-style character consistency.
With your LTX MSR implementation, is there a way to bind specific text tokens or attention conditioning to individual subject references throughout the denoising process?
Could this potentially prevent character attribute swapping, or does LTX's architecture make that kind of token locking impractical?
I kind of stopped making this thing because i finally got v2 of msr and stability with 2 people and settled with this instead of trying to code something myself with ai help.
Also i was going to code my own msr style thing to copy what h3 does, because the h3 system is much better, its built in, and i stopped using h3 because i am an American and i cannot release my stuff on youtube with them. and ltx is much better visually to my eyes.
I previously experimented with something I called VAC Decypher. The idea was to track and lock the tokens or conditioning associated with each character, so their identities, clothing, hair, and physical features wouldn't bleed into other characters.
I was trying to achieve something closer to H3-style character consistency.
With your LTX MSR implementation, is there a way to bind specific text tokens or attention conditioning to individual subject references throughout the denoising process?
Could this potentially prevent character attribute swapping, or does LTX's architecture make that kind of token locking impractical?I kind of stopped making this thing because i finally got v2 of msr and stability with 2 people and settled with this instead of trying to code something myself with ai help.
I’ve tried a similar approach, but the model ended up focusing too heavily on the reference image tokens, leading to motion issues or distorted outputs. That said, I may simply not have tuned the conditioning weights carefully enough.
I think a more promising direction would be to take inspiration from H3: add a semantic understanding layer to LTX, feed it visual tokens extracted from the reference images by a multimodal text encoder, and train this new component separately.
Of course, this is still just a hypothesis. I don’t currently have any spare compute to test it.
Also, another direction would be to apply reinforcement learning (RL) on top of my MSR model. Since many of these visual distortions occur unpredictably, I believe RL could help reduce them and improve the overall quality of the generated videos.
Oh wow, i never made it that far i gave up cuz i got v2 and it rocks. Stoked this works, maybe if my series takes off i can afford a pro 5000 that would help.... I might get into programing for ltx tho..
Oh wow, i never made it that far i gave up cuz i got v2 and it rocks. Stoked this works, maybe if my series takes off i can afford a pro 5000 that would help.... I might get into programing for ltx tho..
Really glad V2 is working well for you! Hope your series takes off — I’d love to see what you’re making with it. And definitely give LTX development a go! There’s still plenty to explore, and I’d be happy to hear about any ideas or experiments you come up with. 🙂
Hey, I wanted to thank you, MSR really save me, i had to ditch h3, and this is what saved my project... Thank you.. I was going crazy with v1 but i know the how it works now, top must be brief but very descriptive to nail it. and the background as well. then it seems to work better... I have very complex characters.....
I forgot to ask, Does vram play a role in precision? having more vram does it improve the ability to have more characters that are stable?
Hey, I have another question about how MSR processes reference images.
I've been experimenting with 3- and 4-panel character reference sheets showing different angles of the same character (front, side, back, and facial close-up).
After improving my TOP character descriptions, I've had much better identity consistency, including a promising three-character scene on my RTX 5070 Ti 16GB.
I'd love to understand how MSR actually interprets these reference sheets.
Multiple views: Does MSR understand that the different panels represent the same character from different angles, or does it process the entire sheet as one composite image?
Panel count: Is a single four-view reference sheet better than supplying separate images for each angle? Does MSR support combining multiple references for one subject effectively?
Resolution: When MSR processes a multi-panel sheet, does it resize the entire image before encoding? Could four panels reduce the effective resolution and quality of each character view?
Backgrounds: Are plain white backgrounds preferable for reference images, or can environmental backgrounds improve conditioning?
Text descriptions: How strongly does the TOP character description influence the reference conditioning? I've noticed that detailed descriptions of facial features, hair, clothing, and body structure dramatically improve character consistency.
Character separation: When using three or more subjects, does physical spacing between characters help prevent identity bleeding or cloning? I've noticed problems when characters enter close together, but better results when they're spatially separated.
My main question is: What would you consider the optimal reference-image format for maintaining character identity across multiple camera angles and scenes, especially with three or more characters?
I'm trying to get the absolute best results possible from your MSR implementation, and understanding how it processes these images would help enormously.
Thanks again for your work!
Hey, I have another question about how MSR processes reference images.
I've been experimenting with 3- and 4-panel character reference sheets showing different angles of the same character (front, side, back, and facial close-up).
After improving my TOP character descriptions, I've had much better identity consistency, including a promising three-character scene on my RTX 5070 Ti 16GB.
I'd love to understand how MSR actually interprets these reference sheets.
Multiple views: Does MSR understand that the different panels represent the same character from different angles, or does it process the entire sheet as one composite image?
Panel count: Is a single four-view reference sheet better than supplying separate images for each angle? Does MSR support combining multiple references for one subject effectively?
Resolution: When MSR processes a multi-panel sheet, does it resize the entire image before encoding? Could four panels reduce the effective resolution and quality of each character view?
Backgrounds: Are plain white backgrounds preferable for reference images, or can environmental backgrounds improve conditioning?
Text descriptions: How strongly does the TOP character description influence the reference conditioning? I've noticed that detailed descriptions of facial features, hair, clothing, and body structure dramatically improve character consistency.
Character separation: When using three or more subjects, does physical spacing between characters help prevent identity bleeding or cloning? I've noticed problems when characters enter close together, but better results when they're spatially separated.
My main question is: What would you consider the optimal reference-image format for maintaining character identity across multiple camera angles and scenes, especially with three or more characters?
I'm trying to get the absolute best results possible from your MSR implementation, and understanding how it processes these images would help enormously.
Thanks again for your work!
Based on my testing:
- MSR can understand multi-angle reference images.
- Even if different angles of the same character are placed in different reference slots, the model can still understand that fairly well.
- Yes, the reference images are resized to match the aspect ratio of your target video. So if you want better facial consistency, I usually recommend including a close-up of the face somewhere in the reference image.
- In my tests, MSR does not handle character references with backgrounds very well. It may even try to copy the reference image too literally, almost like fixing it as the first frame.
- Character descriptions have a very strong influence on the result. If the description is too strong or inaccurate, it can directly change the appearance of the referenced character. You can actually test this by intentionally giving it the wrong description.
- When characters are too close together, especially in scenes where they overlap or interact physically, confusion can happen easily. Even with some spacing, if their outfits are too similar in color or style, identity mixing can still occur. Sometimes the video starts fine, but after a camera cut later in the clip, characters may begin to merge or bleed into each other.
Overall, I’d recommend using a face close-up plus full-body three-view references for each character, along with a clear character description.
I also suggest splitting a scene into smaller parts when possible. For example, instead of generating all characters together in one shot, you can break it into:
- a main confrontation shot for the key characters
- and separate reaction shots for supporting characters
In general, it’s better not to generate everyone together in the same scene unless necessary.
You can also refer to my validation folder for examples.
I still dont know what vram does to msr.... i only have 16, and i can only get two stable characters... so 96 on your has 3 stable characters? or 4? or 5?
I still dont know what vram does to msr.... i only have 16, and i can only get two stable characters... so 96 on your has 3 stable characters? or 4? or 5?
The max count of stable characters on my 6000 is 3
so i spend 16k to have 3. lol ... man 10 maybe i should stay at 16...