Rank Model / vendor Best at Released Configuration Input Output AIdaily Score
01
Balanced across analysis, reasoning, coding, and real repository work
Released 2026-06-09
Configuration max
Input $10
Output $50
80.8
Four-pillar evidence Completeness: 23/23 tasks · 7/7 categories · 4/4 pillars
General 82.33
Science and complex reasoning 92.82
Coding 85.99
Agent / real repository work 62.17
Official page LiveBench source data
02
Strong at complex reasoning and real repository work
Released 2026-07-24
Configuration max
Input $5
Output $25
78.9
Four-pillar evidence Completeness: 23/23 tasks · 7/7 categories · 4/4 pillars
General 75.67
Science and complex reasoning 93.47
Coding 81.45
Agent / real repository work 65.20
Official page LiveBench source data
03
Strong at both complex reasoning and coding
Released 2026-07-09
Configuration max
Input $5
Output $30
78.5
Four-pillar evidence Completeness: 23/23 tasks · 7/7 categories · 4/4 pillars
General 79.79
Science and complex reasoning 93.93
Coding 83.94
Agent / real repository work 56.21
Official page LiveBench source data
04
Strong at understanding requests and handling real repository work
Released To verify
Configuration default
Input $3
Output $15
78.2
Four-pillar evidence Completeness: 23/23 tasks · 7/7 categories · 4/4 pillars
General 78.42
Science and complex reasoning 87.10
Coding 82.47
Agent / real repository work 64.65
Official page LiveBench source data
05
Strong at understanding requests and handling real repository work
Released 2026-07-16
Configuration default
Input $3
Output $15
77.4
Four-pillar evidence Completeness: 23/23 tasks · 7/7 categories · 4/4 pillars
General 78.54
Science and complex reasoning 87.55
Coding 81.45
Agent / real repository work 62.17
Official page LiveBench source data
06
Especially strong at following instructions and everyday analysis
Released 2026-04-23
Configuration xhigh
Input $5
Output $30
77.2
Four-pillar evidence Completeness: 23/23 tasks · 7/7 categories · 4/4 pillars
General 79.89
Science and complex reasoning 92.76
Coding 82.15
Agent / real repository work 53.99
Official page LiveBench source data
07
Especially strong at following instructions and everyday analysis
Released 2026-08-13
Configuration high
Input $0.75
Output $3.75
76.4
Four-pillar evidence Completeness: 23/23 tasks · 7/7 categories · 4/4 pillars
General 77.78
Science and complex reasoning 90.63
Coding 78.89
Agent / real repository work 58.28
Official page LiveBench source data
08
Especially strong at sustained work inside real code repositories
Released 2026-07-19
Configuration max
Input $2
Output $6
76.2
Four-pillar evidence Completeness: 23/23 tasks · 7/7 categories · 4/4 pillars
General 77.39
Science and complex reasoning 89.76
Coding 72.87
Agent / real repository work 64.65
Official page LiveBench source data
09
Especially strong at mathematics, science, and multi-step reasoning
Released 2026-08-12
Configuration default
Input $2
Output $6
75.5
Four-pillar evidence Completeness: 23/23 tasks · 7/7 categories · 4/4 pillars
General 76.47
Science and complex reasoning 91.54
Coding 76.78
Agent / real repository work 57.02
Official page LiveBench source data
10
Especially strong at mathematics, science, and multi-step reasoning
Released 2026-07-09
Configuration max
Input $2.5
Output $15
75.4
Four-pillar evidence Completeness: 23/23 tasks · 7/7 categories · 4/4 pillars
General 75.61
Science and complex reasoning 92.77
Coding 78.25
Agent / real repository work 54.95
Official page LiveBench source data
11
Strong at complex reasoning and real repository work
Released To verify
Configuration xhigh
Input $3
Output $15
75.3
Four-pillar evidence Completeness: 23/23 tasks · 7/7 categories · 4/4 pillars
General 70.19
Science and complex reasoning 90.82
Coding 80.68
Agent / real repository work 59.39
Official page LiveBench source data
12
Strong at both everyday analysis and complex reasoning
Released 2026-03-05
Configuration xhigh
Input $2.5
Output $15
75.0
Four-pillar evidence Completeness: 23/23 tasks · 7/7 categories · 4/4 pillars
General 77.39
Science and complex reasoning 91.13
Coding 77.54
Agent / real repository work 53.84
Official page LiveBench source data
13
Especially strong at sustained work inside real code repositories
Released 2026-08-18
Configuration default
Input $1.4
Output $4.4
75.0
Four-pillar evidence Completeness: 23/23 tasks · 7/7 categories · 4/4 pillars
General 73.13
Science and complex reasoning 86.85
Coding 78.95
Agent / real repository work 60.91
Official page LiveBench source data
14
Strong at both everyday analysis and complex reasoning
Released 2026-08-13
Configuration default
Input $0.435
Output $0.87
74.7
Four-pillar evidence Completeness: 23/23 tasks · 7/7 categories · 4/4 pillars
General 76.34
Science and complex reasoning 90.46
Coding 77.16
Agent / real repository work 54.95
Official page LiveBench source data
15
Especially strong at writing correct code and completing programs
Released 2026-04-16
Configuration xhigh
Input $5
Output $25
74.3
Four-pillar evidence Completeness: 23/23 tasks · 7/7 categories · 4/4 pillars
General 74.31
Science and complex reasoning 90.02
Coding 82.09
Agent / real repository work 50.66
Official page LiveBench source data
16
Especially strong at sustained work inside real code repositories
Released 2026-08-21
Configuration default
Input $0.22
Output $0.66
74.2
Four-pillar evidence Completeness: 23/23 tasks · 7/7 categories · 4/4 pillars
General 76.93
Science and complex reasoning 86.60
Coding 68.20
Agent / real repository work 65.10
Official page LiveBench source data
17
Strong at both complex reasoning and coding
Released 2026-04-16
Configuration max
Input $5
Output $25
74.2
Four-pillar evidence Completeness: 23/23 tasks · 7/7 categories · 4/4 pillars
General 72.58
Science and complex reasoning 91.75
Coding 81.83
Agent / real repository work 50.50
Official page LiveBench source data
18
Especially strong at sustained work inside real code repositories
Released 2026-08-14
Configuration default
Input $0.265
Output $0.65
73.7
Four-pillar evidence Completeness: 23/23 tasks · 7/7 categories · 4/4 pillars
General 74.53
Science and complex reasoning 83.12
Coding 75.69
Agent / real repository work 61.36
Official page LiveBench source data
19
Strong at understanding requests and handling real repository work
Released To verify
Configuration default
Input $2
Output $6
72.5
Four-pillar evidence Completeness: 23/23 tasks · 7/7 categories · 4/4 pillars
General 75.79
Science and complex reasoning 89.00
Coding 68.59
Agent / real repository work 56.46
Official page LiveBench source data
20
Especially strong at writing correct code and completing programs
Released 2025-12-18
Configuration default
Input $1.75
Output $14
72.3
Four-pillar evidence Completeness: 23/23 tasks · 7/7 categories · 4/4 pillars
General 72.78
Science and complex reasoning 83.24
Coding 83.62
Agent / real repository work 49.39
Official page LiveBench source data
21
Especially strong at mathematics, science, and multi-step reasoning
Released 2026-02-05
Configuration high
Input $5
Output $25
72.1
Four-pillar evidence Completeness: 23/23 tasks · 7/7 categories · 4/4 pillars
General 72.16
Science and complex reasoning 89.00
Coding 78.18
Agent / real repository work 48.99
Official page LiveBench source data
22
Especially strong at writing correct code and completing programs
Released 2026-07-09
Configuration max
Input $1
Output $6
72.0
Four-pillar evidence Completeness: 23/23 tasks · 7/7 categories · 4/4 pillars
General 70.24
Science and complex reasoning 86.42
Coding 82.91
Agent / real repository work 48.43
Official page LiveBench source data
23
Especially strong at mathematics, science, and multi-step reasoning
Released 2025-12-11
Configuration high
Input $1.75
Output $14
71.9
Four-pillar evidence Completeness: 23/23 tasks · 7/7 categories · 4/4 pillars
General 73.25
Science and complex reasoning 88.19
Coding 76.07
Agent / real repository work 50.25
Official page LiveBench source data
24
Especially strong at following instructions and everyday analysis
Released 2026-05-19
Configuration high
Input $1.5
Output $9
71.8
Four-pillar evidence Completeness: 23/23 tasks · 7/7 categories · 4/4 pillars
General 75.01
Science and complex reasoning 85.12
Coding 78.18
Agent / real repository work 48.99
Official page LiveBench source data
25
Especially strong at writing correct code and completing programs
Released To verify
Configuration default
Input $1.4
Output $4.4
71.6
Four-pillar evidence Completeness: 23/23 tasks · 7/7 categories · 4/4 pillars
General 70.76
Science and complex reasoning 84.20
Coding 79.65
Agent / real repository work 51.77
Official page LiveBench source data
26
Especially strong at following instructions and everyday analysis
Released 2026-07-31
Configuration default
Input $0.14
Output $0.28
70.8
Four-pillar evidence Completeness: 23/23 tasks · 7/7 categories · 4/4 pillars
General 74.67
Science and complex reasoning 86.71
Coding 74.98
Agent / real repository work 46.77
Official page LiveBench source data
27
Strong at both understanding requests and writing code
Released To verify
Configuration high
Input $1.5
Output $7.5
70.3
Four-pillar evidence Completeness: 23/23 tasks · 7/7 categories · 4/4 pillars
General 74.09
Science and complex reasoning 85.78
Coding 77.86
Agent / real repository work 43.43
Official page LiveBench source data
28
Especially strong at writing correct code and completing programs
Released 2026-02-17
Configuration medium
Input $3
Output $15
70.1
Four-pillar evidence Completeness: 23/23 tasks · 7/7 categories · 4/4 pillars
General 72.42
Science and complex reasoning 85.88
Coding 79.27
Agent / real repository work 42.63
Official page LiveBench source data
29
Especially strong at writing correct code and completing programs
Released 2025-11-01
Configuration high
Input $5
Output $25
69.3
Four-pillar evidence Completeness: 23/23 tasks · 7/7 categories · 4/4 pillars
General 72.75
Science and complex reasoning 85.24
Coding 79.65
Agent / real repository work 39.70
Official page LiveBench source data
30
Especially strong at following instructions and everyday analysis
Released To verify
Configuration max
Input $2.5
Output $7.5
69.3
Four-pillar evidence Completeness: 23/23 tasks · 7/7 categories · 4/4 pillars
General 75.19
Science and complex reasoning 84.29
Coding 74.22
Agent / real repository work 43.59
Official page LiveBench source data
Brand overview
General models Several current models can remain under one vendor, so a whole brand is not judged by a single model.
Especially strong at writing correct code and completing programs
2026-07-09
Availability Public APIStage StableAccess apiPrice $1 / $6Source
Strong at both complex reasoning and coding
2026-07-09
Availability Public APIStage StableAccess apiPrice $5 / $30Source
Especially strong at mathematics, science, and multi-step reasoning
2026-07-09
Availability Public APIStage StableAccess apiPrice $2.5 / $15Source
Especially strong at following instructions and everyday analysis
2026-04-23
Availability Public APIStage StableAccess apiPrice $5 / $30Source
Especially strong at writing correct code and completing programs
2026-03-17
Availability Public APIStage StableAccess apiPrice $0.75 / $4.5Source
Especially strong at mathematics, science, and multi-step reasoning
2026-03-17
Availability Public APIStage StableAccess apiPrice $0.2 / $1.25Source
Strong at both everyday analysis and complex reasoning
2026-03-05
Availability Public APIStage StableAccess apiPrice $2.5 / $15Source
Especially strong at writing correct code and completing programs
2025-12-18
Availability Public APIStage StableAccess apiPrice $1.75 / $14Source
Especially strong at mathematics, science, and multi-step reasoning
2025-12-11
Availability Public APIStage StableAccess apiPrice $1.75 / $14Source
Strong at complex reasoning and real repository work
2026-07-24
Availability Public APIStage StableAccess apiPrice $5 / $25Source
Balanced across analysis, reasoning, coding, and real repository work
2026-06-09
Availability Public APIStage StableAccess apiPrice $10 / $50Source
Especially strong at writing correct code and completing programs
2026-04-16
Availability Public APIStage StableAccess apiPrice $5 / $25Source
Strong at both complex reasoning and coding
2026-04-16
Availability Public APIStage StableAccess apiPrice $5 / $25Source
Especially strong at writing correct code and completing programs
2026-02-17
Availability Public APIStage StableAccess apiPrice $3 / $15Source
Especially strong at mathematics, science, and multi-step reasoning
2026-02-05
Availability Public APIStage StableAccess apiPrice $5 / $25Source
Especially strong at writing correct code and completing programs
2025-11-01
Availability Public APIStage StableAccess apiPrice $5 / $25Source
Strong at complex reasoning and real repository work
To verify
Availability Public APIStage StableAccess apiPrice $3 / $15Source
Especially strong at following instructions and everyday analysis
2026-08-13
Availability Public APIStage StableAccess apiPrice $0.75 / $3.75Source
Especially strong at following instructions and everyday analysis
2026-05-19
Availability Public APIStage StableAccess apiPrice $1.5 / $9Source
Especially strong at writing correct code and completing programs
To verify
Availability Public APIStage StableAccess apiPrice $0.3 / $2.5Source
Strong at both understanding requests and writing code
To verify
Availability Public APIStage StableAccess apiPrice $1.5 / $7.5Source
Not publicly usable now, so it is excluded from the ranking
To verify
Availability Public APIStage PreviewAccess apiPrice $2 / $12Source
Not publicly usable now, so it is excluded from the ranking
2026-08-05
Availability PreviewStage PreviewAccess api, consumerPrice $1.25 / $4.25Source
Not publicly usable now, so it is excluded from the ranking
To verify
Availability PreviewStage PreviewAccess api, consumerPrice $1.25 / $4.25Source
Especially strong at mathematics, science, and multi-step reasoning
2026-08-12
Availability Public APIStage StableAccess apiPrice $2 / $6Source
Strong at both complex reasoning and coding
To verify
Availability Public APIStage StableAccess apiPrice $1.25 / $2.5Source
Strong at understanding requests and handling real repository work
To verify
Availability Public APIStage StableAccess apiPrice $2 / $6Source
Especially strong at sustained work inside real code repositories
To verify
Availability Public APIStage StableAccess apiPrice $1 / $2Source
Best for coding and enterprise agents with European deployment options
To verify
Availability ConsumerStage StableAccess consumerPrice To verifySource
Especially strong at writing correct code and completing programs
2026-04-02
Availability Public APIStage StableAccess apiPrice $0.5 / $3Source
Especially strong at following instructions and everyday analysis
To verify
Availability Public APIStage StableAccess apiPrice $2.5 / $7.5Source
Especially strong at sustained work inside real code repositories
2026-08-14
Availability Open weightsStage StableAccess weightsPrice $0.265 / $0.65Source
Especially strong at sustained work inside real code repositories
2026-07-19
Availability Open weightsStage StableAccess weightsPrice $2 / $6Source
Especially strong at writing correct code and completing programs
To verify
Availability Open weightsStage StableAccess api, weightsPrice $0.6 / $3.6Source
Best for Chinese tasks and agent workflows in Tencent Cloud
To verify
Availability ConsumerStage StableAccess consumerPrice To verifySource
Best for Chinese writing, Q&A, and longer reasoning tasks
2026-05-09
Availability ConsumerStage StableAccess consumerPrice To verifySource
Especially strong at sustained work inside real code repositories
2026-08-21
Availability Public APIStage ExperimentalAccess apiPrice $0.22 / $0.66Source
Strong at both everyday analysis and complex reasoning
2026-08-13
Availability Open weightsStage StableAccess api, weightsPrice $0.435 / $0.87Source
Especially strong at following instructions and everyday analysis
2026-07-31
Availability Open weightsStage StableAccess api, weightsPrice $0.14 / $0.28Source
Especially strong at mathematics, science, and multi-step reasoning
2026-04-24
Availability Open weightsStage StableAccess api, weightsPrice $0.435 / $0.87Source
Balanced across analysis, reasoning, coding, and real repository work
To verify
Availability Open weightsStage StableAccess api, weightsPrice $0.14 / $0.28Source
Strong at understanding requests and handling real repository work
2026-07-16
Availability Open weightsStage StableAccess api, weightsPrice $3 / $15Source
Especially strong at writing correct code and completing programs
To verify
Availability Open weightsStage StableAccess api, weightsPrice $0.95 / $4Source
Strong at coding and sustained work in real repositories
To verify
Availability Open weightsStage StableAccess api, weightsPrice $0.95 / $4Source
Especially strong at following instructions and everyday analysis
To verify
Availability Public APIStage StableAccess apiPrice $0.3 / $1.2Source
Especially strong at sustained work inside real code repositories
2026-08-18
Availability Open weightsStage StableAccess apiPrice $1.4 / $4.4Source
Especially strong at writing correct code and completing programs
To verify
Availability Open weightsStage StableAccess api, weightsPrice $1.4 / $4.4Source
Thinking Machines 1 models
Especially strong at sustained work inside real code repositories
To verify
Availability Open weightsStage StableAccess weightsPrice $1.87 / $4.68Source
Strong at understanding requests and handling real repository work
To verify
Availability Public APIStage StableAccess api, consumerPrice $3 / $15Source
Especially strong at sustained work inside real code repositories
2026-08-20
Availability Public APIStage ExperimentalAccess apiPrice $0 / $0Source
Public picks
General models Only models available to consumers or through a public API are recommended, grouped by practical use.
Overall capability
Balanced across analysis, reasoning, coding, and real repository work
2026-06-09
Availability Public APIStage StableAccess apiPrice $10 / $50Source
Strong at complex reasoning and real repository work
2026-07-24
Availability Public APIStage StableAccess apiPrice $5 / $25Source
Strong at both complex reasoning and coding
2026-07-09
Availability Public APIStage StableAccess apiPrice $5 / $30Source
Coding and real repository work
Balanced across analysis, reasoning, coding, and real repository work
2026-06-09
Availability Public APIStage StableAccess apiPrice $10 / $50Source
Strong at understanding requests and handling real repository work
To verify
Availability Public APIStage StableAccess api, consumerPrice $3 / $15Source
Strong at complex reasoning and real repository work
2026-07-24
Availability Public APIStage StableAccess apiPrice $5 / $25Source
Complex reasoning
Strong at both complex reasoning and coding
2026-07-09
Availability Public APIStage StableAccess apiPrice $5 / $30Source
Strong at complex reasoning and real repository work
2026-07-24
Availability Public APIStage StableAccess apiPrice $5 / $25Source
Balanced across analysis, reasoning, coding, and real repository work
2026-06-09
Availability Public APIStage StableAccess apiPrice $10 / $50Source
Price friendly
Especially strong at following instructions and everyday analysis
2026-08-13
Availability Public APIStage StableAccess apiPrice $0.75 / $3.75Source
Especially strong at mathematics, science, and multi-step reasoning
2026-08-12
Availability Public APIStage StableAccess apiPrice $2 / $6Source
Especially strong at mathematics, science, and multi-step reasoning
2026-07-09
Availability Public APIStage StableAccess apiPrice $2.5 / $15Source
Open weights
Strong at understanding requests and handling real repository work
2026-07-16
Availability Open weightsStage StableAccess api, weightsPrice $3 / $15Source
Especially strong at sustained work inside real code repositories
2026-07-19
Availability Open weightsStage StableAccess weightsPrice $2 / $6Source
Especially strong at sustained work inside real code repositories
2026-08-18
Availability Open weightsStage StableAccess apiPrice $1.4 / $4.4Source
No comparable capability ranking is available yet. These are public specialist models, and AIdaily does not force different modalities into one score.
Best for fast, cost-efficient image generation and editing at scale
2026-02-26
Availability ConsumerStage StableAccess consumerPrice To verifySource
Best for turning detailed instructions into editable finished images
To verify
Availability ConsumerStage StableAccess consumerPrice To verifySource
Best for bolder, more reliable images with stronger personalization
2026-07-24
Availability ConsumerStage StableAccess consumerPrice To verifySource
Best for complex image work that benefits from visual reasoning before generation or editing
2026-02-13
Availability ConsumerStage StableAccess consumerPrice To verifySource
Several current models can remain under one vendor, so a whole brand is not judged by a single model.
Best for turning detailed instructions into editable finished images
To verify
Availability ConsumerStage StableAccess consumerPrice To verifySource
Best for fast, cost-efficient image generation and editing at scale
2026-02-26
Availability ConsumerStage StableAccess consumerPrice To verifySource
Best for complex image work that benefits from visual reasoning before generation or editing
2026-02-13
Availability ConsumerStage StableAccess consumerPrice To verifySource
Best for bolder, more reliable images with stronger personalization
2026-07-24
Availability ConsumerStage StableAccess consumerPrice To verifySource
Only models available to consumers or through a public API are recommended, grouped by practical use.
New and current models
Best for bolder, more reliable images with stronger personalization
2026-07-24
Availability ConsumerStage StableAccess consumerPrice To verifySource
Best for fast, cost-efficient image generation and editing at scale
2026-02-26
Availability ConsumerStage StableAccess consumerPrice To verifySource
Best for complex image work that benefits from visual reasoning before generation or editing
2026-02-13
Availability ConsumerStage StableAccess consumerPrice To verifySource
No comparable capability ranking is available yet. These are public specialist models, and AIdaily does not force different modalities into one score.
Best for storyboard-controlled, coherent multi-shot video with native audio
2026-02-05
Availability ConsumerStage StableAccess consumerPrice To verifySource
Best for video with stable character motion, expressions, and physical effects
2025-10-28
Availability ConsumerStage StableAccess consumerPrice To verifySource
Best for controlling multi-shot clips with image, video, and audio references
2026-02-12
Availability ConsumerStage StableAccess consumerPrice To verifySource
Best for cinematic video with native audio and detailed shot control
To verify
Availability PreviewStage PreviewAccess apiPrice To verifySource
Several current models can remain under one vendor, so a whole brand is not judged by a single model.
Best for cinematic video with native audio and detailed shot control
To verify
Availability PreviewStage PreviewAccess apiPrice To verifySource
Best for controlling multi-shot clips with image, video, and audio references
2026-02-12
Availability ConsumerStage StableAccess consumerPrice To verifySource
Best for video with stable character motion, expressions, and physical effects
2025-10-28
Availability ConsumerStage StableAccess consumerPrice To verifySource
Best for storyboard-controlled, coherent multi-shot video with native audio
2026-02-05
Availability ConsumerStage StableAccess consumerPrice To verifySource
Only models available to consumers or through a public API are recommended, grouped by practical use.
New and current models
Best for controlling multi-shot clips with image, video, and audio references
2026-02-12
Availability ConsumerStage StableAccess consumerPrice To verifySource
Best for storyboard-controlled, coherent multi-shot video with native audio
2026-02-05
Availability ConsumerStage StableAccess consumerPrice To verifySource
Best for video with stable character motion, expressions, and physical effects
2025-10-28
Availability ConsumerStage StableAccess consumerPrice To verifySource
No comparable capability ranking is available yet. These are public specialist models, and AIdaily does not force different modalities into one score.
Best for expressive multilingual voice work
2026-02-02
Availability ConsumerStage StableAccess consumerPrice To verifySource
Best for multi-speaker dialogue and fine voice control
To verify
Availability PreviewStage PreviewAccess apiPrice To verifySource
Best for low-latency voice agents with tool use and audio reasoning
To verify
Availability ConsumerStage StableAccess apiPrice To verifySource
Several current models can remain under one vendor, so a whole brand is not judged by a single model.
Best for low-latency voice agents with tool use and audio reasoning
To verify
Availability ConsumerStage StableAccess apiPrice To verifySource
Best for multi-speaker dialogue and fine voice control
To verify
Availability PreviewStage PreviewAccess apiPrice To verifySource
Best for expressive multilingual voice work
2026-02-02
Availability ConsumerStage StableAccess consumerPrice To verifySource
Only models available to consumers or through a public API are recommended, grouped by practical use.
New and current models
Best for expressive multilingual voice work
2026-02-02
Availability ConsumerStage StableAccess consumerPrice To verifySource
Best for low-latency voice agents with tool use and audio reasoning
To verify
Availability ConsumerStage StableAccess apiPrice To verifySource
Restricted model watchlist These models are not broadly public, so they are excluded from the Top 30 and public picks.
Claude Mythos 5 · Anthropic Available only to approved Project Glasswing partners. It stays on the watchlist and is excluded from ranking and recommendations.
Removed models hunyuan-turbos-20250313 This model has left the current catalog. This notice preserves older links.
ernie-4-5-turbo This model has left the current catalog. This notice preserves older links.
mistral-medium-3 This model has left the current catalog. This notice preserves older links.
gpt-image-1 This model has left the current catalog. This notice preserves older links.
imagen-4 This model has left the current catalog. This notice preserves older links.
seedream-4 This model has left the current catalog. This notice preserves older links.
midjourney-v7 This model has left the current catalog. This notice preserves older links.
sora-2 This model has left the current catalog. This notice preserves older links.
veo-3 This model has left the current catalog. This notice preserves older links.
seedance-2-5 This model has left the current catalog. This notice preserves older links.
minimax-h3 This model has left the current catalog. This notice preserves older links.
kling-2-1 This model has left the current catalog. This notice preserves older links.
gpt-4o-mini-tts This model has left the current catalog. This notice preserves older links.
claude-mythos-5 This model has left the current catalog. This notice preserves older links.
grok-4.6-xhigh This model has left the current catalog. This notice preserves older links.
midjourney-v8-1 This model has left the current catalog. This notice preserves older links.
AIdaily Capability Reference · Composite v1 Each of four pillars carries 25%. The score covers tasks from one LiveBench release only. It does not represent Chinese ability, price, speed, creative preference, or every kind of agent work.
General 25% + science and complex reasoning 25% + coding 25% + agent / real repository work 25%
A small score gap may not be meaningful. A model missing any task is not ranked, and missing values are never set to zero or reweighted.
Data attributed to LiveBench and Abacus.AI. AIdaily re-aggregates the public results into four pillars; this is not an official LiveBench total. The one-line strength is not vendor marketing.
A state link opens the latest list. Long images record the data date visible when shared.
LiveBench LiveBench source data Datasheet Apache 2.0
Today’s AI Digest All issues