To me it feels so weird that people are trying to push these model for online shopping like "here is how this dress/shirt/pants would look on you". But these models will always make the clothes fit your body and show you in flattering light and so on. How the actual garment fits is still as elusive as before these tools
show comments
weird-eye-issue
The meta keywords in the HTML is very interesting. 100+ references to NSFW topics such as hentai, nudes, etc.
show comments
postalcoder
They must have trained on GPT Image 1 outputs. The yellow tint is unmistakable.
It's hilarious how bad image generation is. I just tried this out and I gave it two actual logos, a full design language spec, and three screenshots of the actual application and it spat out three absolutely awful images, none of which had the same logo as the one I sent.
hessammehr
Very cool but the Arabic text in the title image is obviously and hopelessly broken, which is oddly not the case when actually using the model. Could it be that the hero image was not generated by Qwen Image 3.0?
show comments
simonw
> to precisely describe the full 3×3 grid takes a full 3.7k tokens
It's a shame they didn't share that prompt - it would make that demo more convincing.
embedding-shape
Not a single word about when/if they'll actually release the weights for this, or am I missing it somewhere?
show comments
tarcon
I am surprised by the rather bad output. It doesn't achieve qwen image 1 quality in composition or anatomical correctness. Tested on chat.qwen.ai
I am seeing third legs and glowing eyes. It's a Microsoft Lens level of quality and that one was pulled.
sajithdilshan
I truly wish these models were available when I was in University. As a visual learner it would have been much easier for me to understand certain topics with illustrative diagrams rather than reading a wall of text.
show comments
feverzsj
The "piss filter" is still everywhere.
timedude
Zooming in on mobile on that website causes a large white area to obstruct the page. Might wanna look into that.
As for the image model, wow...
zzleeper
Random question, but has there been any improvement in OCR/document understanding in these newer models? Last time I checked (1mo ago) SOTA was still sadly Gemini, unless you wanted to pay $$$ for e.g. Sol
show comments
Mashimo
> Supports up to 4.5k token input, effortlessly generating complex layouts such as newspapers, storyboards, and exam papers.
Impressive.
Btw, what is currently the best model to run locally on a 16GB Vram? Is it Z-Image Turbo?
show comments
gchokov
It failed to create a simple overlay on a map - something ChatGPT had no issues with.
jcattle
What I can not wrap my head around: How are these models trained?
What training mechanism or model architecture provides the glue to go from human text to images?
Don't you need to have millions of really descriptively labelled images?
show comments
Oarch
I assume Van Gogh didn't paint enough hands to train from!
ninjagoo
The examples posted on their launch blog page are quite impressive, especially for fine details, multi-panel/multi-page and text rendering.
But: not open-source/open-weights, and no indication that weights/source will be released either.
pal9000i
How long until we get rid of the AI "plasticness" in portrait kind of generated images?
show comments
geooff_
No pricing table? No benchmarks? This is just marketingslop.
Image token pricing has been fairly steady while text token prices fall, yet image model release discussion seems to be more focused on how beautiful the women the model generates are versus any sort of substantive discussion.
dsrtslnd23
Seems that this will not be open weights?
dzonga
once a.i images took over - a lot of things that were based on images died.
going to a fancy restaurant coz of 'you can take good pics' - dead - a.i can recreate that cheaper.
which means for a certain demographic - dating apps are dead too - since those were largely based on swiping photos.
the premium of in-person / small intimate events has gone up. likewise meeting in person, or doing things with a person live.
this also means imperfection has gone up in value (imperfection is a human quality).
show comments
topheroo
I’d argue that talking about “authentic” AI-generated images is oxymoronic.
bejd
I wonder if they got permission to generate that (admittedly impressive) Berserk image.
lifthrasiir
> In the three examples below, the model accurately renders Japanese, Korean, and Spanish respectively.
And yet the Korean text is not accurate... [1]
[1] E.g. "드레스 컬렉션 dress collection" has vowels ㅔ mixed with ㅐ, "초웜한" should be "초월한 exceeding", "신키한" should be "실키한 silky", "디자언되다" should be "디자인되다 have been designed", "로얼" should be "로열 royal", and so on.
flakiness
It's a bit shocking to see them showing off blatantly disinformation-al examples. See the Japanese manga one. It says "原作 監修 三浦建太郎" which says its original is written by, and it is supervised the by, the famous author (who died a few years ago so who can he possibly supervise?) and "共同制作 白泉社" saying: collaborated by a (famous Japanese) publisher.
I hope these stakeholders have good enough layers.
ndom91
Again not released on huggingface immediately?
Mr_Eri_Atlov
Image generation is the least interesting and most concerning aspect of AI.
I see lots of useful tools made related to translation and code generation, but image generation seems to be primarily used to deceive, harass, or embelish to the point of questioning what the point of a photo is anymore anyway?
show comments
lalith_c
didn’t anthropic accuse them of stealing fable?
treetalker
The red-dress woman's vestigial pinkie toes …
luciana1u
the natural endpoint of this technology is product photos that look better than the actual product, which is going to make unboxing videos the last remaining source of truth on the internet
maxloh
I am curious whether the model requires a font to be installed. Does it also generate the glyphs for the text?
show comments
dhbradshaw
The generated latex pdf!
saltysalt
It will be interesting to compare this to Flux 2.
xiaoyu2006
The blog write-up style is so casual haha.
jdw64
Wow, it displays Korean properly without breaking.
But there are still a lot of typos. Haha, it's good that Korean displays properly, but there are a lot of incorrect sentences
show comments
m3kw9
Impressive, but their woman face generation always use very similar, too perfect, same prettiness faces, it's very obvious.
gpjanik
The real performance is nowhere close to what is presented in the marketing materials, which is pretty annoying. Especially text rendering and accuracy.
Try asking it for a plot of Polish GDP growth over the past 20 years. It's slop.
show comments
mahimai
interesting
rvz
Midjourney already knew that image generation was going to zero. Again yet another reason why the model was never a moat in the first place.
show comments
spwa4
Appears to be closed-weights entirely. No word at all on any weights release.
viridir
This looks impressive!
arslan9063
THIS IS EXACTLY WHAT I WAS LOOKING FOR
Smaily
why I cant submitte new posts here ? My account since 2016
To me it feels so weird that people are trying to push these model for online shopping like "here is how this dress/shirt/pants would look on you". But these models will always make the clothes fit your body and show you in flattering light and so on. How the actual garment fits is still as elusive as before these tools
The meta keywords in the HTML is very interesting. 100+ references to NSFW topics such as hentai, nudes, etc.
They must have trained on GPT Image 1 outputs. The yellow tint is unmistakable.
https://qianwen-res.oss-accelerate.aliyuncs.com/Qwen-Image/i...
https://qianwen-res.oss-accelerate.aliyuncs.com/Qwen-Image/i...
https://qianwen-res.oss-accelerate.aliyuncs.com/Qwen-Image/i...
https://qianwen-res.oss-accelerate.aliyuncs.com/Qwen-Image/i...
It's hilarious how bad image generation is. I just tried this out and I gave it two actual logos, a full design language spec, and three screenshots of the actual application and it spat out three absolutely awful images, none of which had the same logo as the one I sent.
Very cool but the Arabic text in the title image is obviously and hopelessly broken, which is oddly not the case when actually using the model. Could it be that the hero image was not generated by Qwen Image 3.0?
> to precisely describe the full 3×3 grid takes a full 3.7k tokens
It's a shame they didn't share that prompt - it would make that demo more convincing.
Not a single word about when/if they'll actually release the weights for this, or am I missing it somewhere?
I am surprised by the rather bad output. It doesn't achieve qwen image 1 quality in composition or anatomical correctness. Tested on chat.qwen.ai
I am seeing third legs and glowing eyes. It's a Microsoft Lens level of quality and that one was pulled.
I truly wish these models were available when I was in University. As a visual learner it would have been much easier for me to understand certain topics with illustrative diagrams rather than reading a wall of text.
The "piss filter" is still everywhere.
Zooming in on mobile on that website causes a large white area to obstruct the page. Might wanna look into that.
As for the image model, wow...
Random question, but has there been any improvement in OCR/document understanding in these newer models? Last time I checked (1mo ago) SOTA was still sadly Gemini, unless you wanted to pay $$$ for e.g. Sol
> Supports up to 4.5k token input, effortlessly generating complex layouts such as newspapers, storyboards, and exam papers.
Impressive.
Btw, what is currently the best model to run locally on a 16GB Vram? Is it Z-Image Turbo?
It failed to create a simple overlay on a map - something ChatGPT had no issues with.
What I can not wrap my head around: How are these models trained?
What training mechanism or model architecture provides the glue to go from human text to images?
Don't you need to have millions of really descriptively labelled images?
I assume Van Gogh didn't paint enough hands to train from!
The examples posted on their launch blog page are quite impressive, especially for fine details, multi-panel/multi-page and text rendering.
But: not open-source/open-weights, and no indication that weights/source will be released either.
How long until we get rid of the AI "plasticness" in portrait kind of generated images?
No pricing table? No benchmarks? This is just marketingslop.
Image token pricing has been fairly steady while text token prices fall, yet image model release discussion seems to be more focused on how beautiful the women the model generates are versus any sort of substantive discussion.
Seems that this will not be open weights?
once a.i images took over - a lot of things that were based on images died.
going to a fancy restaurant coz of 'you can take good pics' - dead - a.i can recreate that cheaper.
which means for a certain demographic - dating apps are dead too - since those were largely based on swiping photos.
the premium of in-person / small intimate events has gone up. likewise meeting in person, or doing things with a person live.
this also means imperfection has gone up in value (imperfection is a human quality).
I’d argue that talking about “authentic” AI-generated images is oxymoronic.
I wonder if they got permission to generate that (admittedly impressive) Berserk image.
> In the three examples below, the model accurately renders Japanese, Korean, and Spanish respectively.
And yet the Korean text is not accurate... [1]
[1] E.g. "드레스 컬렉션 dress collection" has vowels ㅔ mixed with ㅐ, "초웜한" should be "초월한 exceeding", "신키한" should be "실키한 silky", "디자언되다" should be "디자인되다 have been designed", "로얼" should be "로열 royal", and so on.
It's a bit shocking to see them showing off blatantly disinformation-al examples. See the Japanese manga one. It says "原作 監修 三浦建太郎" which says its original is written by, and it is supervised the by, the famous author (who died a few years ago so who can he possibly supervise?) and "共同制作 白泉社" saying: collaborated by a (famous Japanese) publisher.
I hope these stakeholders have good enough layers.
Again not released on huggingface immediately?
Image generation is the least interesting and most concerning aspect of AI.
I see lots of useful tools made related to translation and code generation, but image generation seems to be primarily used to deceive, harass, or embelish to the point of questioning what the point of a photo is anymore anyway?
didn’t anthropic accuse them of stealing fable?
The red-dress woman's vestigial pinkie toes …
the natural endpoint of this technology is product photos that look better than the actual product, which is going to make unboxing videos the last remaining source of truth on the internet
I am curious whether the model requires a font to be installed. Does it also generate the glyphs for the text?
The generated latex pdf!
It will be interesting to compare this to Flux 2.
The blog write-up style is so casual haha.
Wow, it displays Korean properly without breaking. But there are still a lot of typos. Haha, it's good that Korean displays properly, but there are a lot of incorrect sentences
Impressive, but their woman face generation always use very similar, too perfect, same prettiness faces, it's very obvious.
The real performance is nowhere close to what is presented in the marketing materials, which is pretty annoying. Especially text rendering and accuracy.
Try asking it for a plot of Polish GDP growth over the past 20 years. It's slop.
interesting
Midjourney already knew that image generation was going to zero. Again yet another reason why the model was never a moat in the first place.
Appears to be closed-weights entirely. No word at all on any weights release.
This looks impressive!
THIS IS EXACTLY WHAT I WAS LOOKING FOR
why I cant submitte new posts here ? My account since 2016