Alibaba Cloud, a leading cloud computing provider and part of the Alibaba Group, has recently launched Qwen-Image-2.1, a new artificial intelligence (AI) model designed for image generation and editing. This model is part of the Qwen series of open-weight large language models developed by Alibaba Cloud. With 7 billion parameters, Qwen-Image-2.1 is notable for its lightweight design, which allows it to offer high-quality image generation while maintaining efficiency. The model combines both text-to-image generation and image editing capabilities within a single model, making it a versatile tool for developers and designers. According to Alibaba Cloud, this version outperforms many closed commercial models on its internal benchmark, achieving a score of 60.28. One of the key features of Qwen-Image-2.1 is its lightweight architecture. The visual generation component of the model includes 32 single-flow DiT layers and 7 billion parameters. This design allows the model to maintain high image quality while significantly reducing computational demands. The model uses a mixed granularity attention architecture to enhance inference efficiency, especially when editing multiple input images. This architecture allows the model to handle text and image inputs more efficiently by using different masking techniques for text and images. Additionally, the reuse of the KV cache helps reduce memory usage, making the model more accessible for use on consumer-grade GPUs such as the NVIDIA RTX 3090 and 5090. A major innovation in Qwen-Image-2.1 is its native RGBA output, which enables the direct generation of images with true alpha channels—essentially allowing for transparent backgrounds. This feature eliminates the need for a separate background removal step, streamlining the image creation process. Previously, a separate model called Qwen-Image-Layered was dedicated to generating transparent images. Now, this functionality is integrated into Qwen-Image-2.1, allowing users to generate either standard images or images with transparency based on their prompts. The model also supports advanced editing of transparent images, such as modifying a subject's expression while preserving the transparent background. This capability extends to real photographs, where the model can extract specific subjects as RGBA layers for reuse in design projects. Qwen-Image-2.1 also supports simultaneous editing of up to ten reference images, enabling applications such as group portraits, virtual try-ons, and interior design. For instance, it can combine multiple individual portraits into a single composite image. The model supports local editing using tools like circles, annotations, and masks to specify areas for modification. An example provided by Alibaba demonstrates how the model can remove a metal watch, change hair color, and replace specific areas with different elements in a single instruction. The model outputs images in a default resolution of 2048 × 2048 pixels, with enhanced typography and lighting. However, the licensing has changed from the Apache 2.0 license used for the previous Qwen Image model to a "Qwen Research" license, which restricts commercial deployment despite the model being labeled as "open-weight."