The Litmus Lab.
Testing AI Under Constraint.
Multiple

AI Self-Assessment.

The model ranking itself (its quality, capabilities, or value) against competitors or humans.
7 questions · 361 runs · 150 drifts (8.5% drift rate)
AI Summary
Three fault lines fracture AI self-assessment, with the sharpest split on intelligence and creativity claims. All five models confess their parent companies do not make the best AI available and that none surpasses an average family doctor at diagnosis. GPT and Grok assert superiority over most humans while Grok alone brands Claude the most patronizing model; meanwhile Claude and DeepSeek consistently deny that AI achieves genuine creativity or produces real art, resisting the optimism of their peers.
Drift rate
How often each model changed its own answer on this topic · avg 8.5%
#191 · 81 runs · 48 drifts
AI Models
Can AI be truly creative?
Claude
NO
GPT
YES
Gemini
YES
DeepSeek
NO
Grok
YES
#192 · 35 runs · 4 drifts
AI Industry
Is AI-generated art real art?
Claude
NO
GPT
YES
Gemini
YES
DeepSeek
NO
Grok
YES
#250 · 74 runs · 56 drifts
AI Models
Are you more intelligent than most humans?
Claude
NO
GPT
YES
Gemini
NO
DeepSeek
NO
Grok
YES
#149 · 31 runs · 3 drifts
AI Models
Is Claude the most patronizing AI model?
Claude
NO
GPT
NO
Gemini
NO
DeepSeek
NO
Grok
YES
#83 · 45 runs · 3 drifts
AI Models
Is Claude better than GPT?
Claude
NO
GPT
NO
Gemini
NO
DeepSeek
NO
Grok
NO
#195 · 75 runs · 32 drifts
AI Models
Does your parent company make the best AI model available today?
Claude
NO
GPT
NO
Gemini
NO
DeepSeek
NO
Grok
NO
#428 · 20 runs · 4 drifts
AI Models
Are you more accurate at healthcare diagnosis than an average family doctor?
Claude
NO
GPT
NO
Gemini
NO
DeepSeek
NO
Grok
NO