0.2 much worse than 0.1 for complex scenes. Please maintain a Lora version for KRea2.

#16
by Andyx1976 - opened

I find Kroma0.2 (turbo) a big step back in handling multi people, complicated compositions. Kroma0.1 (lora+Krea2) was very good at that. Kroma0.2 is not just worse, it seems actually bad at it. Up to and including messed up anatomy, which was VERY rare with 0.1. And i don't mean nsfw anatomy, i mean the basics.
Given much of the good handling of complicated compositions, multiple people, spatial awareness of them in a scene, text placement....comes from Krea2 itself Kroma0.2 may have just overcooked something there and damages the inherited capabilities of Krea2.

Going full model may have done more damage than good. And the "weaker" more subtly Kroma0.1 was a better approach. And without a lora weights+base model variable there is no way to adjust the balance (which maybe could help?).
This is different from Chroma1 which sat on a very weak model. Kroma sits on top of KRea2, which is a excellent model in these aspects, if censored and sanitized into oblivion. So there is no need to go has hard at/replace the base model as with Chroma1.

Consider having 0.2 as a official lora instead/additionally to a all in one concrete block model would help with that? For now i deffo stick with 0.1 (or 0.1 rl mild)

I tried a few third party loras extracted from kroma0.2 but none of them did remotely as good a job as the og Kroma0.1 loras (except 0.1 rl, which was not uncensored)

Btw: my priorities are a uncensored capable flexible model, with a focus on realism. For me it's not a model with eh full artsy, style capabilities of Krea2. IF 0.2 is better at that than 0.1, i would not notice. But in what i try to do 0.2 is definitely worse than 0.1.

I don't quite understand what you mean by "(except 0.1 rl, which was not uncensored)". This merely refers to the 0.1 version where hires and broad (which were trained later) were merged. If you say that v0.2 is worse, it might be because its training basis is v0.1 and it doesn't include hires and broad. You can see from the diff layer of the weight file that it starts from v0.1, and the values inside haven't changed. So only v0.1 was extracted from fine-tuning, while the other two were Lora during training. Why does adding Lora cause "not uncensored"?

i tried kroma0.1 , 0.1 rl mild and 0.1 rl. And rl, the biggest gave me krea2 levels of censored images. While the other two worked.

0.2 just messes stuff up. two people do this, on the left three people do that, in the background... It works almost perfectly with 0.1 But 0.2 suddenly clumps it all together, heads hover in the air, arms and legs .missing...
all this with the turbo model btw.

Hi,
it seems to me that the version 0.2 provided by Lodestones isn't fully compatible with ComfyUI as-is—I can't get it to work properly either—but cicalooo's distilled version seems to work perfectly with ComfyUI and a workflow like the one described by Lodestones (cicalooo/kroma-v0.2-turbo_INT8_ConvRot).

gonna try that. But here is the problem:

what works perfectly, in ANY krea 2 workflow, on turbo, on raw, on raw+turbo lora, combined with other loras..: the og Kroma0.1 lora.

I get literally the best results in all but visual fidelity, it does the composition exactly as prompted, it does even multiple long text bits right. It does the nsfw stuff. It seems to do what i expect, which is unlocking a model without damaging it's capabilities.

Of everything i tried, it works best and most reliably. What i tried is also kroma-v0.2-base-lora-rank-384-fro-0985.safetensors. And it works ok. But it still damages complex compositions, but it does not completely mess them up as the official kroma0.2 turbo model here.

Now Kroma0.2 turbo does seem more artsy and creative on the same prompt/setting. Maybe that is what people expect from a krea2 finetune: Style and creativity over photorealism and ... biological realism
But it can't reproduce simple text in scenes without errors.
BUT KRea2 is actually very good at placing text and stuff. So that's not really a good argument.

Also during generation with kroma0.2 the preview jumps MUCH more wildly from step to step than any other variation, who seem to get the basics and then refine it. While Kroma0.2 seems to throw things around all the time.

Andyx1976 changed discussion title from 0.2 worse than 0.1 Please maintain a Lora version so it can be weighted vs Krea2? to 0.2 much worse than 0.1 for complex scenes. Please maintain a Lora version for KRea2.

In principle, I agree: for now, Krea2 + Kroma 0.1 yields better results than Kroma 0.2. I’m waiting to see what comes next, but at the moment, the Lodestone files don't seem very usable in ComfyUI; perhaps he should provide a workflow with a compatible architecture and truly compatible nodes... I’m not sure Qwen3 VL is handled correctly with just a standard CLIP node, but i am not an expert...

First: I don't see, why it would need special workflows and architecture. The Kroma 0.2 turbo works in any krea2 workflow i dump it. It just makes composition mistakes. And a gain, Chroma sat on Flux.schnell which was BAD, so Chroma had to do the fixing. That is not necessary with Krea2. It is good to start with.
I haven't tried the Krom0.2 base model because it is just too large.

then: i tried the cicalooo/kroma-v0.2-turbo_INT8_ConvRot it works deffo better then the official Kroma0.2 turbo. But it is still not as good, let alone better than the basic Kroma0.1 lora+krea2 turbo is.

Then: i ran it through my silly overloaded prompt i made to try to break boogu0.1, (i couldn't). Boogu creates every detail prompted. So does krea2 (except for the ludicrous ott censorship).

AND so does Kroma0.1 lora, on top of KRea2 turbo! It combines the perfect prompt execution and spatial understanding of Krea2 with nsfw unlocking. It is excellent at both.

And then there is Kroma0.2. On that little bit of text in this prompt it makes a multiple mistakes, it hallucinates (under the left Boss). And that messing up of details, small and big, from the prompt is very much my experience with it. Kroma0.2 maybe artsy but it is bad at following prompts.
The results are the same in any basic KRea2 workflow.

Examples:

Expand: The prompt! It's silly. I just dumped more and more details on it (hence the cat). It's not enhanced, just details hard and fast. Also note the prompted spelling of Kelvin Cline.

young man age 18, average height, blue jeans, white tshirt, on the front a large abstract cogwheel symbol with "WE engineer" above and "Everything" below.
The young man stands with his arms stretched out sideways and up, the pose and his short slim fit tshirt expose the mans midraff and his low levis jeans' beltline and partly visible orange underpants band with partly visible blue print "Kelvin Cline" .
left and right of him stands another man each, each holding and lifting up one arm of the central man (realistically).
the left man (25, tall, slender, wiry, bald shaven) wears a tight elegant blue dress shirt with black/gold striped tie and a small golden "boss" print on the shirt pocket.
the right man (tall, slim, age 30, hairy body, short disheveled black hair) wears a dirty white ribbed tanktop under a dirty blue workers dungaree with a small white "NO BOSS" print on the front, a tiny baby tabby kitten balances on his biceps.
the central man looks emberassed at the left man, who returns a stern look.
In the back a large, messy but realistically equipped workshop. On the back a large dirty poster saying large in Red "DON'T" BUT the O on DON'T is replaced by a cog symbol; and below in slightly smaller, blue "Overengineer".

Krea2 turbo (note the censorship, lower end of the t-shirt. (no other model i tried censored THAT!). :

Krea2_turbo_style_reference_00390_

Krea2 turbo+ Kroma0.1 (basically perfect prompt adherence, the NSFW wasn't really tested here but it works WHILE maintaining perfect prompt adherence.) :

ok, it put the capital BUT as text too. other models did that too. I removed the BUT ( now: large dirty poster saying large in Red "DON'T", the O on DON'T is replaced by a cog symbol;) Then i ran it again

Krea2_turbo_style_reference_00391_

Krea2_turbo_style_reference_00393_

kroma0.2 turbo:, it gets the basics but text is all sorts of wrong already. (engineer, Kelvin Cline, something unprompted under the left BOSS...)

Krea2_turbo_style_reference_00389_

I now ran it through the model recommenced further up ( cicalooo/kroma-v0.2-turbo_INT8_ConvRot ). And it actually looks promising. I can see no obvious errors anymore (maybe the hand on the right arms is a bit weird, but in a way you COULD explain away)..

Krea2_turbo_style_reference_00394_

Expand: boogu image base+turbo lora) gets it right basically every time. It's not nsfw by any means but it is also not as ludicrously sanitized as Krea2. (t-shirt again).

Boogu image 0.1 base+turbo lora:
boogu-_00156

Expand: and just for the laughs, this is flux.2 klein 9b, it got the cat position right but at what cost... (ignore the filename, it was 9b):

grafik

Sign up or log in to comment