AI Model Leaderboard

Which AI model actually performs best? Not marketing claims. Not synthetic benchmarks. Every text and image generated on OneAIWorld is scored by an independent AI reviewer — these are the live results.

Run Your Own Comparison
🏆
gemma-4-26b-moe
Top overall — 9.44/10
gpt-oss-120b
Fastest — 8.4s avg
🎨
gemma-4-26b-moe
Most creative — 8.36/10
📊
24 models
201 reviewed generations

Last 90 days of generations. Text models need 5+ and image models 3+ reviewed generations to be ranked — small samples can't distort the board.

Text Models

Scored 0–10 on accuracy, relevance, coherence, creativity, and language quality, plus the reviewer's overall rating — that's the ranking. Strongest At shows where each model beats the field average, so you can pick a model for your kind of task.

Model Overall Accuracy Relevance Coherence Creativity Language Quality Avg Speed Reviews Strongest At
🥇 gemma-4-26b-moe 9.44 9.5 10.0 9.57 8.36 9.79 41.4s 14 Creativity
🥈 gemma-4-31b 9.38 9.42 10.0 9.33 8.33 9.83 51.6s 12 Creativity
🥉 claude-haiku-4-5 9.15 9.33 10.0 9.25 7.67 9.5 18.9s 12 Accuracy
4 claude-opus-4-6 9.08 9.12 9.5 9.38 7.62 9.75 46.4s 8 Coherence
5 deepseek-r1 9.07 9.08 9.67 9.08 7.83 9.67 24.3s 12 Creativity
6 gpt-4-1-mini 8.92 9.2 9.9 9.1 6.9 9.5 9.5s 10 Accuracy
7 gpt-5-4 8.91 9.0 9.78 9.11 7.11 9.56 9.8s 9 Coherence
8 llama-4-scout 8.66 8.79 9.64 8.79 6.79 9.29 13.0s 14 Relevance
9 deepseek-ai/DeepSeek-R1-0528 8.63 8.67 9.67 8.67 7.17 9.0 33.6s 6 Relevance
10 Qwen/Qwen2.5-72B-Instruct 8.60 8.86 9.71 8.71 6.57 9.14 19.3s 7 Relevance
11 mistral-medium 8.60 8.5 9.38 8.88 7.0 9.25 36.6s 8 Coherence
12 Qwen/Qwen3-235B-A22B-Instruct-2507 8.60 8.5 9.5 9.0 6.83 9.17 15.6s 6 Coherence
13 deepseek-ai/DeepSeek-V3 8.60 8.75 9.62 8.62 7.0 9.0 24.8s 8 Relevance
14 grok-4-1-fast-reasoning 8.47 8.44 9.33 8.0 7.67 8.89 15.9s 9 Creativity
15 phi-4 8.42 8.56 9.0 8.67 6.67 9.22 30.0s 9 Language Quality
16 openai/gpt-4o-mini 8.42 8.5 9.33 8.67 6.25 9.33 8.7s 12 Language Quality
17 zai-org/GLM-4.6 8.17 7.83 8.83 8.33 6.83 9.0 20.5s 6 Creativity
18 meta-llama/Llama-4-Scout-17B-16E-Instruct 8.16 8.2 9.0 8.4 6.0 9.2 14.2s 5 Language Quality
19 gpt-oss-120b 8.16 8.11 9.0 7.89 6.78 9.0 8.4s ⚡ fastest 9 Creativity
20 openai/gpt-oss-120b 8.04 8.2 8.6 8.4 6.0 9.0 27.6s 5 Language Quality

Image Models

Scored 0–10 on accuracy, relevance, coherence, creativity, and image quality, plus the reviewer's overall rating. Image models rank with 3+ reviewed generations — image volume runs lower than text, and the sample size is shown per row.

Model Overall Accuracy Relevance Coherence Creativity Image Quality Avg Speed Reviews Strongest At
🥇 gpt-image-1-mini 9.05 9.25 9.75 9.0 8.0 9.25 32.8s 4 Accuracy
🥈 mai-image-2 8.88 9.0 9.6 9.0 7.6 9.2 26.9s 5 Accuracy
🥉 imagen3 8.63 8.33 8.83 9.0 7.5 9.5 10.3s 6 Image Quality
4 flux-dev 8.40 7.4 8.8 8.8 7.4 9.6 10.1s 5 Image Quality

Reliability

Quality only matters if the answer arrives. Success rate of generation jobs by AI platform on our infrastructure.

PlatformJobsSuccess Rate
Azure AI 6 100.0%
Black Forest Labs 6 100.0%
Google 9 100.0%
OpenAI 10 80.0%
Microsoft 8 62.5%

How the Leaderboard Works

✍️

Real generations

Users run real prompts through real models — no synthetic test sets, no cherry-picking.

🤖

Independent review

Every output is scored 0–10 by an independent AI evaluator on five quality dimensions plus an overall rating.

📊

Aggregated per model

Scores, latency, and success rates are averaged per model. Text models need 5+ reviews to rank, image models 3+.

🔄

Always current

Rankings shift with every new generation — this is live real-world output, not a one-time benchmark run.

🔒 Privacy: only aggregate statistics are ever published — never user prompts, and never generated content.

Strongest At compares each model against the field average, so it highlights genuine standout ability rather than dimensions every model aces. For deeper matchup analysis, read our AI model comparisons and guides.

Last updated August 15, 2026.

Which model wins for your prompt?

Type it once — see every model's answer side by side, scored.

Compare Text Models Compare Image Models