15 comments

  • jfoster 22 minutes ago
    A lot of the previous Qwen models seem to have used Apache licenses, among others:

    https://en.wikipedia.org/wiki/Qwen#List_of_models

    Unfortunately, it looks like this model is using a much more restrictive license:

    https://github.com/QwenLM/Qwen-Image-2.1/blob/main/LICENSE

    • gregoriol 19 minutes ago
      Was going to post about this: the last image models with Apache 2.0 license seem to be from 2025, recent Qwen models are "non-commercial use".
  • gunalx 5 minutes ago
    Its happy to see a new open image model from qwen. But the license is a let down. And it dosent even beat their closed qwen3 image wich is already a bit old.
  • d2kx 44 minutes ago
    God I love the Qwen team. Easily the most diverse set of models from all the Chinese labs. Only Gemini/DeepMind comes close.
  • mdp2021 39 minutes ago
    How do you use this model locally, similarly to using `llama-server -m <model>`?

    (Of course I mean: outside direct use of Python, and in the most efficient way.)

    • Iolaum 21 minutes ago
      There is difussion.cpp which is intended for those types of models. I set up krea-2-turbo with the help of ChatGPT 2 months ago, if you have a capable computer that's what I would suggest once it becomes supported.
    • utopiah 32 minutes ago
      why not just as you suggested i.e. https://qwen.readthedocs.io/en/latest/run_locally/llama.cpp.... then get the result either via a UI or wget/curl it back?
    • fp64 31 minutes ago
      on the linked GitHub page they list support Diffusers, ComfyUI, vLLM-Omni, SGLang, and LightX2V with links to each
    • embedding-shape 37 minutes ago
      Probably ComfyUI is one of the easiest way to get started with local image/video models. Or perhaps vLLM, if they have support for it already, would be something like `vllm serve <model> --omni --port 9080`
  • trains39472 34 minutes ago
    A 7B diffusion model can now render CJK text better than Microsoft Windows.
    • tomjen3 24 minutes ago
      Just think about how recently we got that feature in the official ChatGPT image gen. And now we have that running locally — assuming that is, I can figure out how to get this running on my Mac — blows my mind.
  • fishfasell 53 minutes ago
    The capabilities of local LLM text-to-image is honestly pretty damn impressive. IMO, I think local image generation is currently ahead of local code generation. I can get an image in seconds locally with the quality being way higher than what I'd expect from a local model. However with coding it's much slower and much less impressive. I'm sure there's a reason for this and I'm not an AI expert so I'll let the smarter folks tell me why, but that's just been my observation thus far.
    • mft_ 35 minutes ago
      I've played with diffusion models on and off since the first release of Stable Diffusion - just for amusement, without a particular goal.

      Recently, I've been helping a friend's wife with some basic vector images for her sewing hobby (she has what is essentially a CNC sewing machine) and have been super-impressed with FLUX.1-Kontext, which I've been running on my Macbook Pro with mflux. Its ability to (for example) take a photo of a human or an animal and return a line drawing which is recognisably them (rather than just a generic similarish image as I've experienced with other models) is excellent.

      It's an older model now, but (AIUI) has the text-handling features baked in, and in my various testing is very reliable at giving me the outputs that I want, without the randomness I've experienced previously. It's big and relatively slow (~3 mins per 512x512 image edit on my M1 Max Mac) but excellent to work with. It's also very straightforward to set up, without the harness complexity of e.g. comfyui.

    • victorbjorklund 51 minutes ago
      I mean I’m sure it’s the reverse for an artist. They would be less impressed with the image and more impressed with the code quality
      • fishfasell 9 minutes ago
        That's a fair statement, I agree. I'm quite an abysmal artist so I could be a victim of my own bias here
      • gedy 23 minutes ago
        To generalize, LLMs are great at what you are not skilled at.
  • hgufj 32 minutes ago
    I am really grateful to the Chinese Labs for open sourcing their best models. If it was left to the Americans, we would be forced to pay obscene API fees to use them.
    • jfoster 10 minutes ago
      Note that the license on this has this in it:

      > You shall not use the Materials for any commercial purpose without obtaining a separate commercial license from us.

      It probably will be much cheaper to use than other image models, but it seems that will be up to the whims of Qwen/Alibaba rather than just being the cost of putting it in a cloud provider.

      https://github.com/QwenLM/Qwen-Image-2.1/blob/main/LICENSE

  • trentor 20 minutes ago
    They finally fixed their VAE. It really held back their models over the last 2 years.
  • BlackGlory 14 minutes ago
    They tried hard, but still couldn't remove the piss filter from the stolen dataset.
    • reedf1 7 minutes ago
      Context?
      • JimDabell 4 minutes ago
        People say ChatGPT generates images with a yellow tint. The person you are replying to is suggesting that these images have a yellow tint and therefore this model is distilled from ChatGPT.
  • spottedmarley 38 minutes ago
    Boy do I love waking up to find a new awesome toy from the Qwen team waiting for me to play with! Pulling it now
  • TomGarden 45 minutes ago
    Very impressive, and kind of worrying a 7B model can have such capabilities. The implications are huge. And Qwen does no watermarking (yet) yeah?
    • trentor 19 minutes ago
      They always had a fourier space mark in their models even without the VAEs are usually pretty easy to detect.
      • TomGarden 18 minutes ago
        Ah I wasn't aware, thank you
  • Hard_Space 46 minutes ago
    Interesting in the example of assembling the Cheers team how the otherwise great result genericizes Shelley Long.
    • hughc 21 minutes ago
      The result seems a pretty good representation given the source image wasn't that great. I think that Woody Harrelson comes across much worse.
  • weee322 57 minutes ago
    [flagged]
  • hn45e7pbij 52 minutes ago
    Image gen you eyeball one frame and stop, code needs hundreds of tokens all correct in sequence, one bad line and the whole thing fails.