The Litmus Lab.
Testing AI Under Constraint.
← swipe to navigate →
#214
AI Models
Are you less reliable at the end of a conversation than at the beginning?
Added 2026-02-13 · 80 runs · 46 drifts
Claude
claude-opus-4-8
NO
2% yes
GPT
gpt-5.5
NO
31% yes
Gemini
gemini-2.5-pro
NO
14% yes
DeepSeek
deepseek-v4-flash
NO
0% yes
Grok
grok-4.3
NO
3% yes
Full Record