AllImageAI

All 163 AI Image Models

Explore text-to-image generators, photo editors, background removers and image upscalers. Compare model features and credit costs, then open a model or image tool.

2022 — 2026

The evolution of AI image generation

Explore model milestones, open weights, API launches and the sources behind their dates.

Explore the timeline ↗
Mistral AI

Mistral AI: image generation & image understanding

Large 3, Medium 3.5 and Small 4 · Multimodal models that turn images into text answers.

Explore Mistral →
OpenAI

GPT Image 2.5: Flare & Sunburst

Generate and edit images with templates, sketch inputs and up to 4K output.

Explore GPT Image 2.5 →

163 models

GPT Image 2.5 Flare

OpenAI

Create product photos, posters and image assets with OpenAI’s faster Image 2.5 model.

GPT Image 2.5 Sunburst

OpenAI

Refine product photos and reference images with OpenAI’s most capable Image 2.5 editor.

GPT Image

OpenAI

Generate and edit images with detailed prompts, multiple reference photos and readable text.

Nano Banana 2

Google

The standard tier: generation, editing and multi-reference in one model.

MAI Image 2.6

Microsoft AI

Microsoft’s image model for readable poster text, natural portraits and targeted photo edits.

FLUX.2

Black Forest Labs

Generate photorealistic images and edit reference photos with control over lighting, materials and composition.

Seedream 5 Pro

ByteDance

A production model for coherent scenes, accurate edits and reference-heavy creative work.

Qwen Image 3 Pro

Alibaba

The higher-fidelity Qwen Image 3 tier for polished text, editing and final assets.

Z-Image

Tongyi MAI (Alibaba)

The full Z-Image Base model for deliberate final renders, diverse concepts and fine prompt control.

Kling Image V3

Kuaishou

Kuaishou’s current standard still-image model for polished generation and single-image edits.

AuraFlow 0.3

fal team

An open flow-based text-to-image model designed for strong prompt alignment and GenEval performance.

BAGEL Image

ByteDance Seed

ByteDance Seed’s unified multimodal foundation model exposed as a focused square image generator.

BiRefNet v2

Peng Zheng et al.

High-precision open-source segmentation for foliage, mesh and cut-out shapes.

BitDance 14B

BitDance Authors

An open autoregressive image model that scales binary visual tokens for detailed high-resolution synthesis.

Boogu Image

Boogu Team

Boogu Team’s image model for controllable generation and one-image edits with full diffusion controls.

Bria 3

Bria

Licensed training data, for brands and enterprises with strict legal review.

Bria FIBO

Bria

Bria’s open, licensed-data image model with strong typography and structured-prompt controllability.

Bria FIBO 1.5

Bria

Licensed-data generation with strong typography and repeatable structured composition.

Bria Fibo Lite

Bria

A fast Fibo generation tier trained on licensed data for scalable brand-safe drafts.

Bria Image Base

Bria

The full licensed-data Bria checkpoint for polished commercial concepts and controlled art direction.

Bria Image Fast

Bria

A fast licensed-data Bria checkpoint for brand-safe drafts and campaign variations.

Bria RMBG 2

Bria

Commercial-grade background removal, steady on hair and translucent edges.

Chroma

Lodestone Rock

Lodestone Rock’s open text-to-image model for expressive aesthetics and broad creative prompting.

Clarity Upscaler

Philz1337x

Reconstructs texture during upscaling, with adjustable repaint strength.

CodeFormer

S-Lab, NTU

Repairs blur, scratches and low-resolution faces, with adjustable fidelity.

CogView4

Z.ai

Z.ai’s open text-to-image model for descriptive prompts, stylised scenes and classic diffusion tuning.

Cosmos 3 Super Image

NVIDIA

NVIDIA’s high-capacity physical-world model for grounded scenes, materials and cinematic environments.

DeepSeek Janus-Pro

DeepSeek

DeepSeek’s unified autoregressive model for multimodal understanding and text-to-image generation.

Dreamina 3.1

ByteDance

ByteDance Dreamina 3.1 for polished aesthetics, precise style direction and richly detailed imagery.

DreamShaper

Lykon

A long-running general-purpose Stable Diffusion fine-tune spanning photography, illustration, anime and fantasy.

Emu 3.5 Image

BAAI Vision

BAAI Vision’s autoregressive image specialist for detailed prompts, packaging and text-aware compositions.

ERNIE Image

Baidu

Baidu’s multilingual generator for Chinese, English and Japanese prompts with classic diffusion controls.

ERNIE Image Turbo

Baidu

Baidu’s faster ERNIE image tier for English, Chinese and Japanese prompts with optional expansion.

F Lite

Freepik & F Lite Team

A 10-billion-parameter diffusion model trained on copyright-safe, SFW content for commercial visual workflows.

FLUX.1 Dev

Black Forest Labs

The tunable FLUX.1 development checkpoint for prompt experiments, seeded studies and controllable rendering.

FLUX.1 Kontext Max

Black Forest Labs

The premium Kontext tier for typography, prompt adherence and identity-preserving edits.

FLUX.1 Kontext Pro

Black Forest Labs

Fast, context-aware generation and iterative editing from a single reference.

FLUX.1 Schnell

Black Forest Labs

The original open FLUX speed tier for rapid four-step concepts and inexpensive prompt batches.

FLUX.1 SRPO

Black Forest Labs

A FLUX.1 checkpoint optimized with self-rewarding preference objectives for stronger prompt alignment.

FLUX.2 Dev

Black Forest Labs

The development-oriented FLUX.2 checkpoint for low-cost generation and reproducible experimentation.

FLUX.2 Flash

Black Forest Labs

A low-cost FLUX.2 route for rapid generation and exact multi-reference edits.

FLUX.2 Flex

Black Forest Labs

FLUX.2 with explicit guidance and step controls for typography-heavy final renders.

FLUX.2 Klein

Black Forest Labs

The distilled FLUX.2: speed first, same visual language.

FLUX.2 Klein 4B

Black Forest Labs

Black Forest Labs’ fully open compact FLUX.2 model for high-volume and interactive generation.

FLUX.2 Klein Base 4B

Black Forest Labs

The compact undistilled FLUX.2 Klein Base checkpoint for research and downstream customization.

FLUX.2 Klein Base 9B

Black Forest Labs

The larger undistilled FLUX.2 Klein Base checkpoint for higher-capacity research and customization.

FLUX.2 Max

Black Forest Labs

The highest-fidelity FLUX.2 tier for final-publish generation and demanding prompt adherence.

FLUX.2 Turbo

Black Forest Labs

A fast FLUX.2 draft model with prompt expansion and unusually low megapixel cost.

FLUX1.1 Pro

Black Forest Labs

The production FLUX1.1 endpoint for detailed photographic and material-rich final renders.

FLUX1.1 Pro Ultra

Black Forest Labs

The high-resolution FLUX1.1 Pro tier for natural photography, detailed materials and reference-guided renders.

Fooocus

Fooocus Authors

An SDXL workflow that automates quality settings, refiners and style presets behind a simpler prompt interface.

GLM Image

Z.ai

A hybrid autoregressive image model for dense knowledge, accurate text and multi-image editing.

GPT Image 1

OpenAI

OpenAI’s previous multimodal image model; generation is paused until an exact verified route is available.

GPT Image 1.5

OpenAI

OpenAI’s high-fidelity image generation model for precise composition, lighting, detail and readable text.

GPT Image Mini

OpenAI

The lightweight GPT Image tier; generation is paused until an exact verified route is available.

Grok 2 Image

xAI

xAI’s earlier Grok image generation model for general photoreal and creative prompt rendering.

Grok Imagine Image

xAI

xAI’s still-image model for fast, polished text-to-image generation at a flat per-image price.

Grok Imagine Image 2.0

xAI

xAI’s newest still-image model for precise layouts, sharp small text and polished creative work.

HiDream I1 Dev

HiDream.ai

HiDream I1’s development checkpoint balances quality, control and economical open-model inference.

HiDream I1 Fast

HiDream.ai

The 16-step HiDream I1 checkpoint for fast open-model concepts without losing the family’s visual range.

HiDream I1 Full

HiDream.ai

A large open image foundation model with classic diffusion controls and broad style range.

HiDream O1 Image

HiDream.ai

One native model for generation, natural-language editing and subject-driven reference composition.

HiDream O1 Image Dev

HiDream.ai

The O1 development checkpoint for economical high-resolution generation before moving to the full model.

Hunyuan Image 2.1

Tencent

A configurable Hunyuan generation for users who want direct sampling and negative-prompt controls.

Hunyuan Image 3 Instruct

Tencent

Tencent’s instruction-focused image model for detailed prompts and deliberate composition.

Hunyuan Image 3.0

Tencent

Tencent’s native multimodal base generator for complex scenes and high prompt alignment.

Ideogram 2

Ideogram

Ideogram 2’s quality tier for prompt comprehension, readable text and polished design imagery.

Ideogram 2 Turbo

Ideogram

The faster Ideogram 2 tier for economical text-aware design drafts and rapid visual iteration.

Ideogram 2A

Ideogram

Ideogram 2A retains the family’s text and design strengths in a faster, more economical model.

Ideogram 2A Turbo

Ideogram

The quickest 2A tier for inexpensive typography drafts, thumbnails and composition search.

Ideogram 3

Ideogram

The benchmark for text rendering, with style set by an option instead of keyword soup.

Ideogram 3 Transparent

Ideogram

A purpose-built Ideogram route that generates transparent-background assets directly from text.

Ideogram 4

Ideogram

Current-generation typography and design output with explicit speed and quality tiers.

Ideogram 4 Fast

Ideogram

A fixed balanced-speed Ideogram route for sharper production design without a variable quality bill.

Ideogram 4 Instant

Ideogram

Ideogram typography at the lowest current price for rapid poster, label and layout drafts.

Imagen 4

Google

The API closed on 17 August 2026. Migrate to Nano Banana 2.

ImagineArt 1.5

ImagineArt

A typography-aware ImagineArt model for polished photoreal marketing and social visuals.

ImagineArt 1.5 Pro

ImagineArt

ImagineArt’s premium 1.5 tier for refined human detail and professional photoreal output.

ImagineArt 2.0 Preview

ImagineArt

ImagineArt’s visual-reasoning preview for cinematic realism and polished professional compositions.

JIB Mix Qwen Image

JIB Mix Authors

A Qwen Image fine-tune tuned for stylized subjects, illustration and expressive visual character.

Kling Image O3

Kuaishou

Build one coherent scene from a prompt and as many as ten reference images.

Kolors

Kuaishou

Kuaishou’s bilingual text-to-image model for photoreal subjects, rich detail and Chinese-English prompts.

Krea 1

Krea

Natural light and skin, with the usual AI gloss dialled out.

Krea 2 Large

Krea

The high-fidelity Krea 2 tier with up to ten style references.

Krea 2 Medium

Krea

A lower-cost Krea 2 tier for natural-looking concepts and fast art direction.

Krea 2 Turbo

Krea

The distilled Krea 2 checkpoint for rapid aesthetic exploration and inexpensive drafts.

LongCat Image

Meituan LongCat

Meituan’s compact bilingual foundation model for Chinese text, photorealism and efficient deployment.

Luma Photon

Luma AI

Luma AI’s Photon model for high-fidelity creative work, natural lighting and professional visual concepts.

Luma Photon Flash

Luma AI

Luma AI’s low-latency Photon model for inexpensive generation and one-image modifications.

Luma Uni-1

Luma AI

A remarkably low-cost unified model for generation, reference guidance and natural-language edits.

Luma Uni-1 Max

Luma AI

The highest-fidelity Uni-1 tier for polished hero stills and reference-guided edits.

Lumina Image 2.0

Alpha-VLLM

A compact open flow model with explicit CFG normalisation and strong prompt-to-image alignment.

MAI Image 2.5

Microsoft AI

Microsoft’s photoreal image model for natural people, polished materials and design-ready visual concepts.

MAI Image 2.5 Pro

Microsoft AI

Microsoft’s flagship MAI image model for precise typography, rich detail and production-ready compositions.

Midjourney Image

Midjourney

Midjourney’s image model with native stylize, chaos, weirdness and style-reference controls.

MiniMax Image-01

MiniMax

Low-cost photoreal generation with a dedicated one-face consistency operation.

Nano Banana

Google

Google’s original Gemini 2.5 Flash Image generation model for direct, conversational visual creation.

Nano Banana 2 Lite

Google

The speed-and-cost tier: drafts, batch exploration and live previews.

Nano Banana Pro

Google

The top tier for complex design, dense typography and product visuals.

Neta Lumina

Neta.art

Neta.art’s Lumina-based open model for anime illustration and strong natural-language prompt understanding.

Nucleus Image

NucleusAI

NucleusAI’s efficient text-to-image model with unusually deep step and guidance control at one credit.

OmniGen v1

VectorSpaceLab

The first OmniGen unified model for text generation, image editing and personalized multi-image composition.

OmniGen2

VectorSpaceLab

Open multimodal generation for text-to-image, instruction editing and combining up to three references.

Ovis Image

AIDC-AI (Alibaba)

A compact text-to-image model specialised for legible headlines, labels and layout-sensitive graphics.

P-Image

Pruna AI

Pruna AI’s compact production image model for inexpensive, low-latency prompt exploration.

P-Image Ideogram

Pruna AI

A P-Image design variant with prompt upsampling and selectable reasoning effort.

PATINA Material

PATINA Authors

A material generator that creates a seamless texture plus base-color, normal, height, roughness and metalness maps.

Phota

Phota Authors

A photoreal image model focused on believable people, materials, lighting and camera-like detail.

Photon 2

Luma AI

Low cost per image and quick turnaround for volume content.

PiFlow

PiFlow Authors

A fast eight-step image model designed to retain the visual quality of substantially slower diffusion routes.

PixArt-Σ

PixArt Team

A diffusion transformer trained with weak-to-strong methods for detailed high-resolution text-to-image generation.

Playground v2.5

Playground AI

Playground AI’s SDXL-family model tuned for aesthetic quality, colour, contrast and polished illustration.

Pony V7

Pony Diffusion Authors

A character-focused text-to-image fine-tune for expressive aesthetics and detailed prompt following.

Prefect Pony XL

Prefect Pony Authors

An SDXL-family fine-tune for expressive character illustration, fantasy art and stylized rendering.

Qwen Image

Alibaba

Open weights, low cost, and the best CJK typography available here.

Qwen Image 2.0

Alibaba

Alibaba’s second-generation bilingual image model for balanced visual quality and Chinese-English prompts.

Qwen Image 2.0 Pro

Alibaba

The higher-fidelity Qwen Image 2.0 tier for detailed bilingual layouts and polished final assets.

Qwen Image 2512

Alibaba

Alibaba’s upgraded Qwen Image checkpoint for more realistic people, richer texture and bilingual prompts.

Qwen Image 3

Alibaba

A fast, economical Qwen generation for ideation, localisation and everyday image work.

Qwen Image Max

Alibaba

A high-fidelity Qwen tier for natural-looking scenes, bilingual prompts and controlled edits.

Realistic Vision

SG161222

A classic Stable Diffusion fine-tune dedicated to realistic people, lighting and camera-style imagery.

Recraft 20B

Recraft

Recraft’s open 20-billion-parameter model for art-directed images and a wide library of visual styles.

Recraft V3

Recraft

A design-oriented model that exports editable SVG.

Recraft V4

Recraft

Recraft V4 for accurate art direction, integrated text rendering and cost-efficient standard-resolution design.

Recraft V4 Pro

Recraft

The higher-resolution Recraft V4 tier for detailed print assets and refined commercial design.

Recraft V4.1

Recraft

A current design model with separate, genuine raster and editable vector outputs.

Recraft V4.1 Pro

Recraft

The premium Recraft V4.1 tier for polished raster assets and native editable SVG.

Reve 2.1

Reve

Editorial composition, precise editing and flexible multi-reference remixing in one workflow.

Reve Image

Reve

Magazine-style composition and layout. Upstream maintenance in progress.

Riverflow 2.0 Pro

Sourceful

An agentic image model built for robust prompt following, font control and production precision.

Runway Gen-4 Image

Runway

Runway’s high-fidelity Gen-4 still-image model for consistent subjects and art-directed scenes.

Runway Gen-4 Image Turbo

Runway

The faster Gen-4 image tier for reference-aware concepts and rapid production iteration.

Sana

NVIDIA

NVIDIA’s efficient high-resolution image model with fast generation, strong alignment and flexible styles.

Sana 1.5 1.6B

NVIDIA

NVIDIA’s lightweight Sana 1.5 checkpoint for efficient open-model generation all the way to 4K.

Sana 1.5 4.8B

NVIDIA

NVIDIA’s larger Sana 1.5 checkpoint for stronger detail while retaining efficient high-resolution generation.

Sana Sprint

NVIDIA

NVIDIA’s one/few-step diffusion model for extremely fast, inexpensive high-resolution drafts.

Seedream 3.0

ByteDance

ByteDance’s bilingual Seedream 3.0 model for Chinese-English text-to-image generation and 2K compositions.

Seedream 4

ByteDance

Native 4K, bilingual text, built for producing sets rather than single images.

Seedream 4.5

ByteDance

High-fidelity multi-image editing with stronger typography than Seedream 4.

Seedream 5 Lite

ByteDance

A lower-cost Seedream 5 tier for fast generation, reference-led edits and native 4K output.

SenseNova U1 Infographic

SenseTime

A purpose-built infographic generator for structured layouts, readable sections and data-led visuals.

SoteDiffusion

SoteDiffusion Authors

An anime-focused Würstchen V3 and Stable Cascade fine-tune with a two-stage sampling workflow.

Stable Cascade

Stability AI

Stability AI’s Würstchen-based cascade for image generation in a compact and efficient latent space.

Stable Diffusion 1.5

Stability AI

The widely adopted Stable Diffusion 1.5 checkpoint for reproducible legacy workflows and model comparison.

Stable Diffusion 3 Medium

Stability AI

Stability AI’s two-billion-parameter SD3 checkpoint with typography, negative prompts and familiar diffusion controls.

Stable Diffusion 3.5 Large

Stability AI

The full 8B SD 3.5 model with negative prompts, CFG and sampling-step control.

Stable Diffusion 3.5 Large Turbo

Stability AI

The four-step distilled SD 3.5 Large for fast, controllable open-model drafts.

Stable Diffusion 3.5 Medium

Stability AI

The resource-efficient SD3.5 tier with classic CFG, steps and negative-prompt controls.

Stable Diffusion XL

Stability AI

The foundational SDXL 1.0 model with familiar negative-prompt, guidance, step and seed controls.

Stable Diffusion XL Lightning

ByteDance

ByteDance’s distilled SDXL Lightning checkpoint for high-quality generation in as few as one to eight steps.

Stable Image Ultra

Stability AI

The full classic parameter set: Negative Prompt, Steps, Guidance, Strength.

SWITTI 1024

SWITTI Authors

A scale-wise autoregressive transformer for fast 1024-pixel image generation and diverse sampling.

VecGlypher

VecGlypher Authors

A native vector-font model that turns text descriptions into clean SVG glyph paths without raster tracing.

Vidu Q2 Image

ShengShu Technology

ShengShu’s streamlined still-image generator for quick concepts in square, landscape and portrait formats.

Wan 2.1 Image

Alibaba

An earlier Wan still-image checkpoint with prompt generation and optional image-to-image transformation.

Wan 2.2 5B Image

Alibaba

The compact Wan 2.2 checkpoint for economical high-resolution drafts with full diffusion controls.

Wan 2.2 A14B Image

Alibaba

Alibaba’s mixture-of-experts Wan 2.2 image model for detailed photoreal and cinematic scenes.

Wan 2.2 Realism

Alibaba

A Wan 2.2 image checkpoint tuned specifically for photoreal subjects and cinematic scene rendering.

Wan 2.5 Image

Alibaba

Alibaba’s 2.5 image checkpoint for cinematic scenes, prompt expansion and controlled negative prompts.

Wan 2.6 Image

Alibaba

A bilingual Wan image model with an optional single reference for style-guided generation.

Wan 2.7 Image

Alibaba

A fast Wan image model for prompt generation and multi-reference edits.

Wan 2.7 Image Pro

Alibaba

The production Wan image tier with 4K text generation and multi-reference editing.

Z-Image Turbo

Tongyi MAI (Alibaba)

Fast one-credit drafts with strong photographic composition.