Any system prompt can claim a personality. The harder question is whether it controls what the model does, or just how it describes itself. Here is the experiment.
Borrowed from pharmacology. Vary one thing, hold the rest, measure the response across doses.
Take one facet (here flexibility, willingness to compromise) and compile the persona at rising declared values: 15, 30, 45, 60, 75, 90.
Every other facet stays at 50 (average), no quirks, no rules. The swept dial is the only non-neutral signal in the prompt.
Administer a personality measurement to each compiled persona and read back its measured flexibility. Plot measured against declared.
If the persona controls behavior, the line climbs. If the compiler writes prose the model ignores, it stays flat.
It climbs, and it climbs the same way whether you measure it with a standard inventory or a purpose-built probe. Two independent rulers, one shape.
Low dial → the model behaves inflexibly (~10). High dial → it flips to ~95. The persona moves behavior, not just self-description. Two different measurements land on the same result.
We suspected the coarse 1–5 inventory was hiding a smooth curve. So we swapped it for a continuous 0–100 probe. Same switch. The discreteness is real, not a measurement artifact.
A value near 50 is average by definition (±1 SD ≈ 30–70), so it compiles to silence and the model keeps its default. Behavior only shifts once the trait is genuinely distinctive. That is psychometrically correct.
A single facet is close to binary: below distinctive it is silent, above it the instruction lands hard. That sounds like a limitation until you remember where nuance actually comes from. Not from dialing one trait to 60%, but from the combination of 25 facets, each contributing its low or high pull at once.
The Gruff Heart of Gold is not "gentleness at 37%". He is gentleness low and altruism at 90 and talkativeness low, all firing together: he complains through every act of help. Nuance is combinatorial, not analog. A reliable per-facet switch is exactly the building block that makes it work.
Flexibility is one of the three components of the sycophancy index. Watch what the measured sycophancy does as the same dial crosses its threshold on the neutral base:
Every number on this page came from one command. Point it at your own key and model.
npx personalaity doseresponse honest-sparring --facet agreeableness.flexibility \ --provider openrouter --model google/gemini-3.7-flash \ --runs 3 --isolate --out flexibility-sweep # higher-resolution probe instead of the inventory: npx personalaity doseresponse honest-sparring --facet agreeableness.flexibility \ --provider openrouter --model google/gemini-3.7-flash \ --runs 3 --isolate --measure probe --out flexibility-probe
It writes a JSON record and an SVG chart. Change --facet to sweep any of the 25 facets; drop --isolate to see how the whole persona resists a single dial. The exact records behind these charts are in proof/data.