How Storming measures, and where its limits are
This page is for readers with a background in psychometrics. It describes the instruments, the scoring and the rules in detail, and closes with what has not yet been established.
Instrument catalogue
| Key | Instrument | Items | Purpose | Licence |
|---|---|---|---|---|
| ipip-neo-120 | IPIP-NEO-120 | 120 | Big Five and 30 facets | public domain |
| ipip-50 | IPIP Big-Five Factor Markers | 50 | Five main dimensions only, for a shorter start | public domain |
| mini-ipip | Mini-IPIP | 20 | Re-test and comparison with the earlier profile | public domain |
| ipip-ipc | IPIP Interpersonal Circumplex | 32 | Dominance × warmth (in preparation, not yet available) | public domain |
| climate-6 | Own team-climate instrument | 6 | Team climate, aggregate only | self-authored, own copyright |
| culture-24 | Own culture instrument | 24×2 | The team's culture map | self-authored, own copyright |
| change-pulse | Own single-question pulse | 1 | Where people stand in a change, self-rated | self-authored |
Every instrument in the catalogue may be used commercially without a licence fee or attribution. We left out a values instrument because it did not meet that condition.
Why no types
Personality traits are continuously distributed, and most people sit somewhere in the middle. A type system draws a line at one point, so two people who answer almost identically end up in different categories as soon as they fall just either side of it. The category then quickly becomes an expectation of the person.
Types are popular because a label is easier to remember than a percentile, but research supports them far less than dimensional models such as the Big Five. Storming therefore describes where someone sits on each dimension and never assigns anyone a type.
Scoring
Answers → validity check → raw score → norm table → percentile with confidence interval. The check catches incomplete questionnaires and long runs of identical answers. If it fails, no profile is computed and the person takes the questionnaire again. The lead is not told.
Norms
Scores are compared with the full reference sample, and the percentile says so: “relative to the full reference sample”. Norms split by group are not used. We do not offer sex-specific norms at all, because a statement like “high for a woman” serves no legitimate purpose at work.
The measurement-error rule
Person A is at the 57th percentile, with an interval of 46 to 68. Person B is at 48, interval 37 to 59. The two scores are nine points apart, but the intervals overlap by thirteen. There is no difference the data can support, and Storming does not show one.
The rule is implemented in the scoring itself rather than in the display, so it applies to every view and to the pair rules as well.
The eight starter rules
The table is generated from the same files the application evaluates. Every rule is versioned and available in full at /rules/, with its rationale in German and English.
| Rule | Condition | What you would see | Proposed agreement |
|---|---|---|---|
| agreeableness-gap v1.0.0 |
|Δ A| ≥ 25 pts, intervals disjoint | Directness reads as aggression one way, restraint as evasiveness the other — neither reading is what was meant. | An explicitly agreed "I see this differently" phrasing counts as safe. Nothing is read between the lines. |
| both-high-assertive v1.0.0 |
both E3 ≥ p70 | person A and person B compete for airtime; decisions get reopened after the meeting — by whoever spoke second. | Rotate who runs the agenda. Every decision is written down and owned by one named person. |
| both-low-assertive v1.0.0 |
both E3 ≤ p30 | Decisions sit unclaimed because neither person A nor person B takes the call — both wait politely. | Name a decider per topic in advance — arbitrarily if need be. Arbitrarily decided beats undecided. |
| conflict-norm-clash v1.0.0 |
|Δ culture open_conflict| ≥ 30 | Each believes their own conflict norm is the team's norm: person A and person B experience the same exchange as normal and as a violation. | Surface both readings in the team retro — the clarification belongs in the team, not in the pair. |
| openness-gap v1.0.0 |
|Δ O4| ≥ 25 pts, intervals disjoint | "Let's stick with what works" lands as obstruction, "let's just try it" as recklessness — both are meant as care. | Separate the "whether" question from the "how" question: decide together whether — then plan separately how. |
| orderliness-gap v1.1.0 |
|Δ C2| ≥ 25 pts, intervals disjoint | What one person considers done, the other considers open — both have redone work the other thought was finished. | Before work starts, write down what "done" means per artefact — ticket, migration, runbook. |
| pace-clash v1.0.0 |
|Δ culture decision_speed| ≥ 30 | One ships to learn; the other wants the analysis first — each experiences the other as reckless or as slow. | Time-box the analysis explicitly and agree the reversibility test: what would be hard to undo? Only there does analysis come first. |
| withdrawal-asymmetry v1.0.0 |
A N4 ≥ p70, B A ≤ p30 | Directness from person B lands considerably harder on person A than intended; person A raises things directly less and less. | person B states the intent before the critique. person A gets 24 hours to respond in writing instead of having to react on the spot. |
More rules will be added once they have been justified and reviewed. The rationale is always in the file itself.
Limitations
What is not yet established or still has to be tested:
- We developed the culture instrument ourselves. Before its results can be relied on, it needs validation with at least 30 respondents and a factor analysis. The same applies to the climate instrument.
- The norms come from a large international online sample of self-selected participants. Whether the percentiles fit Swiss teams, in engineering for example, just as well has not been tested.
- Personality scores describe tendencies. They can explain why something is harder for someone, but they cannot predict what that person will do in a given situation.
- The eight pair rules are reasoned assumptions that have not yet been confirmed empirically. Whether the agreements they lead to hold up in practice is recorded and feeds into revising the rules.