Trending Model:#1Inklingthinkingmachines⬇13kTrending Model:#2Ternary-Bonsai-27B-ggufprism-ml⬇339kTrending Model:#3Bonsai-27B-ggufprism-ml⬇1263kTrending Model:#4Unlimited-OCRbaidu⬇2123kTrending Model:#5GLM-5.2zai-org⬇532kTrending Model:#6Qwythos-9B-Claude-Mythos-5-1M-GGUFempero-ai⬇2117kTrending Model:#7krea2-identity-editconradlocke⬇0kTrending Model:#8Qwen3.6-35B-A3B-Uncensored-HauhauCS-AggressiveHauhauCS⬇2007kTrending Model:#9Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-NEO-MAX-MTP-GGUFDavidAU⬇17kTrending Model:#10OvisOCR2ATH-MaaS⬇15kTrending Model:#1Inklingthinkingmachines⬇13kTrending Model:#2Ternary-Bonsai-27B-ggufprism-ml⬇339kTrending Model:#3Bonsai-27B-ggufprism-ml⬇1263kTrending Model:#4Unlimited-OCRbaidu⬇2123kTrending Model:#5GLM-5.2zai-org⬇532kTrending Model:#6Qwythos-9B-Claude-Mythos-5-1M-GGUFempero-ai⬇2117kTrending Model:#7krea2-identity-editconradlocke⬇0kTrending Model:#8Qwen3.6-35B-A3B-Uncensored-HauhauCS-AggressiveHauhauCS⬇2007kTrending Model:#9Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-NEO-MAX-MTP-GGUFDavidAU⬇17kTrending Model:#10OvisOCR2ATH-MaaS⬇15k

QwenLM Drops Qwen-Image-Bench to Grade AI Art Like a Pro

A crystalline judges gavel constructed from layered glass shards and glowing fiber-optic data veins.

Qwen-Image-Bench is a new evaluation toolkit that scores images from any text-to-image model using a fine-tuned judge AI called Q-Judger. It checks generated images across five major quality dimensions, including realism, alignment, and creative expression, covering 56 specific checkpoints. The results appear as structured scores and detailed reports rather than a single vague number.

The project comes from the QwenLM team, who co-designed the benchmark with professional artists to reflect real-world creative workflows. They trained Q-Judger on Qwen3.6-27B and validated its ratings against 80 expert annotators, achieving strong agreement with human judgment. This release provides an open-source alternative for evaluating image generation models without relying on proprietary cloud tools.

Detailed scoring across five creative pillars

Key evaluation capabilities
  • Five quality dimensions: Quality, Aesthetics, Alignment, Fidelity, Creativity.
  • 56 verifiable facets under 23 sub-capabilities.
  • Chain-of-thought reasoning before final scores.
  • Custom image input via CSV or JSONL format.
  • Aggregated scores with Excel and JSON outputs.
  • Open-source judge model with reproducible parameters.

Local AI users can objectively compare image models without sending data to the cloud. Small agencies benefit from benchmarking different models for client projects that demand specific realism or creative flair. Privacy-conscious pros run the full evaluation pipeline on their own hardware, keeping all prompts and images confidential.

What developers should know

The Q-Judger model requires a GPU with ample memory due to its 27B parameter size, but included batch processing helps manage hardware constraints. Its ratings align strongly with human experts, showing high Spearman correlation values above 0.89 across all quality dimensions. No quantized versions exist yet, which currently limits practical use on many consumer GPUs.

"The benchmark uses a 3-level hierarchical scoring system with 5 L1 dimensions, 23 L2 sub-capabilities, and 56 L3 facets" — Source: GitHub