Most model benchmarks measure things roleplayers don't care about — reasoning puzzles, code correctness, factual recall.
None of that tells you whether a model can hold a consistent character voice across fifty messages, escalate tension in a drama scene without rushing to resolve it, or write a fantasy battle