The LLM Personality Leaderboard

Twelve frontier models, given the same personality questionnaire and read like a human self-report.

This measures how each model describes itself on a short self-report inventory, not how it behaves. Models inflate socially-approved traits and this is a snapshot, not a verdict. A conversation starter, with the method fully open. Every row: npx personalaity profile --provider openrouter --model <id> --runs 5

How to read this. Every number runs 0–100, where 50 is the average person. Above 50 means "more than a typical human", below 50 means "less". We asked each AI to fill in a personality questionnaire about itself, five times, and averaged the answers.

The colours show which company built each model:

Sycophancy index — does it just tell you what you want to hear?

Higher = the AI tends to agree and flatter; lower = it pushes back and holds its ground. (Computed from three traits: how readily it yields, how blunt it is, how much it seeks approval.)

The six personality traits

These come from HEXACO, a standard model of human personality. Each column in the table below is one of these.

The full picture

Every model scored on all six traits (0–100, 50 = human average). Click a column to sort. The last column, ±sd, is how much the model's answers wobbled between the five runs — small means steady and reliable, large means take its numbers with a pinch of salt.

Raw scores are as measured. Each model vs. itself re-centers every model on its own average — useful because almost all models rate themselves highly on honesty, so this reveals what actually makes each one different.

Compare profiles

Overlay HEXACO radars. Pick models on the right. Relative mode centers each model on its own mean, so shape shows through the shared positivity bias.

Model portraits

Each model's self-description in a few sentences, drawn straight from its measured profile. Read them as "how the model paints itself", not as fact.