embedding-shape

> We found that the model's modulation weights (~40% of the total parameters) could be pruned and replaced with a functionally equivalent lookup table, dramatically shrinking the memory footprint with no loss in output quality.

Is this a common approach to reducing weights with "no loss in output quality", assuming this is true? Seems almost too simple to work. If this is doable, would this be applicable to LLMs as well?

Neat with native frame-to-frame generation, but wonder how easy it is to "link" together clips at the intersection, typically the models kind of lose the "momentum" across these stiches, being able to merge things with frame-to-frame between clips might help with this it feels like.

show comments
vblanco

Im running this on my 4070ti super (16 gb vram), and it takes 10 minutes for a 10-seconds 480p video. but the results are spectacular.

show comments
sheesdev

The mouse render is surprisingly good. Several of those clips stood out to be a pretty big leap in terms of current SOTA models.

The only one that looks "off" is the beverage ad video during the can opening clip, it still has that "AI smoothening" effect. Good thing this can be done pretty well using traditional rendering.

I feel like for a good while now we'll transition into a process that uses traditional "close-up" rendering/shots + AI generated wide-shots or quick cuts.

Exciting, but also troubling. This being open-weights is a massive win for the community though.

show comments
Mashimo

> The result gives a total memory footprint reduced by 66%, from 123.6 GB in full precision to 42.5 GB with the smallest models variants. Combining this with our dynamic VRAM offloading enables a next-generation 2K video model to run locally on a GPU like the RTX 3060.

Pretty cool.

But assuming you have a 16GB 3060, how long would it take to generate a 15 second clip?

show comments
fodkodrasz

On one hand: impressive. On the other hands aesthetically it all looks painfully bland and generic.

show comments
mihau

Has anyone tried running this on Mac device? (e.g. Mac Studio Ultra)

show comments
storus

Reference-to-video mode seems like all that was missing to enable completely independent cinematography as right now one couldn't stitch different scenes together properly without altering substantial portions of the scene.

SV_BubbleTime

I saw the samples people have posted. Immediately deleted LTX2 and WAN folders. Those are completely worthless now.

There is some debate on the license for those in the US, UK, EU, plus… no comment other than whew those samples though!

show comments
satvikpendem

I've said it before and I'll say it again, human directors are still valuable, as they use AI video editing tools to generate the shots they want and put them together in a cohesive way. Previously they might've used film and actors but if they can just prompt the AI (or create workflows as seen with ComfyUI) then they arrange them together just like how an EDM producer doesn't actually play the instruments but instead the creativity is in the arrangement.

I suspect it'll be quite a while until AI gets a good enough aesthetic sense to do this, as even with static HTML websites humans can easily see that it's AI slop.

show comments
hnlqpx99l9

Learned something, upvoted

nfnmema

Any tutorial for me to learn how to use

rvz

Hollywood and the film industry on red alert. Too bad.

This is AGI.

show comments