How this test is built
This page explains the result from the inside: what the four numbers are counted from, where the 108 questions came from, where the confidence thresholds sit, and what we honestly do not know yet.
1What the test measures and what it does not
The test reports four numbers, and all four are counted from the same answers. Your type is how your attention is built: where it goes on its own, without you deciding, and which question you are solving all the time without noticing. It is not a set of qualities and not a list of strengths. Two people of the same type can be unlike each other in almost everything except that build. Your wing is the shade one of your two neighbors lends you, and it shows in behavior more than in motive. Your instinctual subtype is where most of your energy is aimed: at keeping yourself provided for, at your standing among people, or at one close bond. It changes the outward pattern of a type so much that it is usually the reason somebody fails to recognize themselves in the description. Your state is a state, not a grade. It says which mode you are in now, and it moves over a lifetime and sometimes over a day. What the test does not do: it does not promise that you will recognize yourself. Sometimes the description lands at once and sometimes it does not, and the second does not mean the count is wrong. It also does not measure how well your life is going, does not predict what you will do, and does not tell you who to be.
2Where the 108 questions came from
All 108 questions were written for this test. Not one of them is taken from somebody else's questionnaire: a borrowed question drags a borrowed model along with it, and what it measures here is something nobody can check. Twelve questions for each of the nine types. Inside each set of twelve it is decided in advance how many questions belong to which mode and which instinct, and that distribution is checked by machine before every build: if it is broken, the build does not pass. Each language version is written from the same construct card rather than translated from another language. That costs roughly twice as much and is done for one reason: translating a question changes what the question measures. What a question may not do: it carries one statement rather than two stitched together, it has no double negatives, and it contains no term from the theory. You should not be able to work out which scale you are feeding.
I usually run a decision past several people first.
A real question from the bank, word for word. The type it belongs to is not named: that part is closed, see below.
3How the result is counted
The main difference from most personality tests is worth saying plainly: we do not compare you against a database of other people. It is usually done the other way round. Your answers are put next to the answers of thousands of others, your position relative to the average is read off, and the result comes out of that. That method has an unpleasant property: your result depends on who took the test before you. Change the sample and you change the type. We count differently. Your answers are brought to your own spread: what matters is not whether you gave something a four, but what you rated higher than everything else of yours. Somebody generous with high ratings and somebody who hoards them come out comparable, because each is compared with themselves. From which follows the part we like best: your result will not change when our database grows. It is counted from your answers and from nothing else. Confidence is counted the same way. We look at how consistently you answered the questions that belong together, and at how far the first candidate pulled ahead of the second. If it barely pulled ahead, we say so rather than rounding to one answer.
4Why the state has no number
The classic model has nine levels and they are numbered. We show three ranges and give no numbers. The reason is arithmetic, not taste. To tell nine steps apart reliably a test needs a precision that a 108-question form does not have and cannot have: it would take something like a thousand. At our length about three steps are honestly distinguishable, and we show three. A number would look more precise and would be invented, and showing nine while telling three apart means selling a precision that is not there. And one thing matters more than the arithmetic: a range shows which way the balance tipped, not which rung you are standing on. The upper questions and the lower ones are not the two ends of one ruler. The upper ones ask what your own strength costs you; the lower ones ask what marks you off from the types next to you when things get hard. Those are separate quantities, so a person can be full of energy and in the grip at the same time, and the middle range does not mean halfway. It means neither side is ahead. The names of the ranges are chosen on purpose too: with room to spare, with neither ahead, in the grip. You will not find the words healthy and unhealthy here, common as they are in the literature. A range describes not a person but the state they are in today, and a label that sounds like a verdict has already gone on to describe the person.
5How often we are wrong
The honest answer: we do not know yet, and we are not going to pretend otherwise. We have not checked a single live result. What we have is a calculation on a model. We described how people answer questions like these, built sixty thousand imaginary people out of that description with their types known in advance, and ran their answers through the same count that produces your result. Then we compared what the count said with what had been put in. The difference between that and a real check is simple: we counted not on people but on a model of people that we built ourselves. If the model is unlike the living, the number has nothing to do with the living. On that calculation, roughly one named type in sixteen turns out to be the wrong one. The wing and the instinct come out similar, about six percent each, and if somebody takes the wing top-up the share drops to about four. The caveat without which these numbers cannot be read: they are counted only among the people we made a statement about at all. Where we show two close candidates instead of one there is no percentage, because there is no claim either. The same view shows why the grades of confidence are needed. Where we name a type confidently, two runs by the same person agree in about 97 cases out of a hundred. Where we do not, agreement falls to almost chance. It is one property seen from two sides: confidence separates what we know from what we do not. The calculation will be replaced with a measurement once we have enough people who have taken the test twice. How exactly is written below.
6What happens when we are not sure
A soft result is not a refusal and not a breakdown. It is the same measurement with its limit named honestly. When the first candidate has not pulled far enough ahead of the second, we show both and name what exactly the person is choosing between. Rounding to one answer would be easier and would look more confident, but it would no longer be our estimate, it would be our decision made for you. The top-up does not exist everywhere, and the difference matters. For the wing it works: what gets in the way there is measurement noise, which is to say we simply did not have enough questions to separate two neighbors, and a dozen more really does separate them. For the instinct there is no top-up, and not because we could not be bothered to build one. For many people the first and second instinct really are close, rather than close for want of data. A top-up in that case runs not into the precision of the measurement but into there being nothing to measure: a difference that does not exist will not appear from new questions. So where two instincts come out almost even, we show both and suggest reading both descriptions rather than taking another test.
7What is closed and why
Almost everything is open. The formulas for all four outputs, the way confidence is counted and the actual thresholds, how the questions are distributed across modes and instincts, where the weights came from, and the list of what we tried and threw away along the way. One thing is closed: which question belongs to which type, and with it the discrimination codes inside the bank. There is exactly one reason and it is worth naming plainly. Knowing which questions map to which type, you can get any result you like. Everything else is open for the opposite reason: it lets you check us rather than fool the test. The line runs exactly there and nowhere else: closed is what could be used to fix the answer, open is what explains the method.
8What this test cannot do
Three things this test does not do, and all three are stated as facts rather than as small print. First. This is not a medical or psychological diagnosis, and it's not a substitute for professional help. It's here to help you understand yourself, and that's all it claims to do. None of the four numbers describes a state of health or replaces a conversation with a professional. Second. The test is not fit for selection. Not for hiring, not for appraisals, not for handing out roles, and not for any decision about a person that somebody else makes. The result is counted from self-report, which is to say from what a person said about themselves on one particular day, and turning that into grounds for somebody else's decision is not fair to them. Third, and this is the least obvious. An instrument agreeing with itself is not the same as being right. A test that steadily measures the wrong thing will show beautiful repeatability: it will be wrong in the same way every time and will look reliable doing it. Everything said above about 97 cases out of a hundred and about retaking belongs to agreement, not to hitting the mark. Checking that what we measure is the Enneagram rather than something next to it is separate work, and it is still ahead of us. We say this for a specific reason: we very nearly took reliability for accuracy ourselves.
9How this will be checked
The mechanism is simple and it ships with the first release rather than sitting in the plans. Some of the people who take the test will get an invitation to take it a second time, free, and will see both of their results side by side. Three things in that arrangement are deliberate. The invitation is random. That matters more than it looks: invite only the people who doubted their result and you will gather plenty of data with an invisible error in it, because the ones who doubted differ from everybody else in exactly the way we are interested in. The gap before the second attempt varies, from a couple of weeks to about a year, and it is assigned at random too. With a fixed gap, memory of the earlier answers and real change in the person stay permanently indistinguishable. Both results are shown to the person in full, including the disagreement between them if there is one. We consider about two and a half thousand pairs enough. When they are gathered, the calculation above will be replaced with a measurement, and a date will stand here.
10Where the model came from
The model goes back to Oscar Ichazo's enneagram of fixations, was carried into the psychology of character by Claudio Naranjo, and was developed by Don Richard Riso and Russ Hudson, with whom the levels of development appeared. The 27 instinctual subtypes were systematized by Naranjo.
We do not use anybody else's questionnaires: this test is not RHETI and not iEQ9, the project is not connected with The Enneagram Institute or with Integrative9, and RHETI and iEQ9 are the trademarks of their owners. All 108 questions were written by us. The list below is the origin of the model and its scientific context, not the source of the questions.
- Hook, J. N., Hall, T. W., Davis, D. E., Van Tongeren, D. R., Conner, M. (2021)The Enneagram: A systematic review of the literature and directions for future researchJournal of Clinical Psychology, 77(4), 865-883A systematic review of Enneagram research: what has been tested and what has not.doi:10.1002/jclp.23097
- Bland, A. M. (2010)The Enneagram: A Review of the Empirical and Transformational LiteratureThe Journal of Humanistic Counseling, Education and Development, 49(1), 16-31A review of the empirical literature of the previous generation.doi:10.1002/j.2161-1939.2010.tb00084.x
- Newgent, R. A., Parr, P. E., Newman, I., Higgins, K. K. (2004)The Riso-Hudson Enneagram Type Indicator: Estimates of Reliability and ValidityMeasurement and Evaluation in Counseling and Development, 36(4), 226-237Estimates of reliability and validity for the Riso-Hudson questionnaire.doi:10.1080/07481756.2004.11909744
- Wagner, J. P., Walker, R. E. (1983)Reliability and validity study of a Sufi personality typology: The enneagramJournal of Clinical Psychology, 39(5), 712-717One of the first psychometric papers on the typology.doi:10.1002/1097-4679(198309)39:5<712::AID-JCLP2270390511>3.0.CO;2-3
- Sutton, A., Allinson, C., Williams, H. (2013)Personality type and work-related outcomes: An exploratory application of the Enneagram modelEuropean Management Journal, 31(3), 234-249The Enneagram set against the Big Five on work outcomes.doi:10.1016/j.emj.2012.12.004
- Daniels, D., Saracino, T., Fraley, M., Christian, J., Pardo, S. (2018)Advancing Ego Development in Adulthood Through Study of the Enneagram System of PersonalityJournal of Adult Development, 25(4), 229-241Work from the Daniels group on adult development through study of the model.doi:10.1007/s10804-018-9289-x
- Alexander, M., Schnipke, B. (2020)The Enneagram: A Primer for Psychiatry ResidentsAmerican Journal of Psychiatry Residents' Journal, 15(3), 2-5A short introduction to the model for psychiatry residents.doi:10.1176/appi.ajp-rj.2020.150301
- Goldberg, L. R., Johnson, J. A., Eber, H. W., Hogan, R., Ashton, M. C., Cloninger, C. R., Gough, H. G. (2006)The international personality item pool and the future of public-domain personality measuresJournal of Research in Personality, 40(1), 84-96The canonical paper on public-domain item pools. Our discipline for writing items follows that tradition; the items themselves are ours.doi:10.1016/j.jrp.2005.08.007
- Meade, A. W. (2004)Psychometric problems and issues involved with creating and using ipsative measures for selectionJournal of Occupational and Organizational Psychology, 77(4), 531-551Why ipsative scales are treacherous. Taken into account in how the scoring is built.doi:10.1348/0963179042596504
- Riso, D. R., Hudson, R. (1996)Personality Types: Using the Enneagram for Self-DiscoveryHoughton Mifflin, ISBN 978-0-395-79867-6The levels of development are described here, and the wings were systematized in the same line of work.
- Riso, D. R., Hudson, R. (1999)The Wisdom of the EnneagramBantam Books, ISBN 978-0-553-37820-1The best known presentation of the Riso and Hudson model.
- Naranjo, C. (1994)Character and Neurosis: An Integrative ViewGateways Books and Tapes, ISBN 978-0-89556-066-7Naranjo carried the model into the psychology of character and developed the 27 instinctual subtypes.
- Palmer, H. (1988)The Enneagram: Understanding Yourself and the Others in Your LifeHarper & Row, ISBN 978-0-06-250673-3The Palmer line: attention as the key to type.
- Daniels, D., Price, V. (2009)The Essential Enneagram: The Definitive Personality Test and Self-Discovery GuideHarperOne, ISBN 978-0-06-171316-3The book that carries the Stanford SEDI questionnaire and its validation data.
- The Oscar Ichazo FoundationArica SchoolOscar Ichazo's school. His enneagram of fixations is where the whole modern typology begins.arica.org
- 特定非営利活動法人 日本エニアグラム学会日本エニアグラム学会The Japanese Enneagram Association. Its materials were used to check the Japanese terminology of this site.www.enneagram.ne.jp
- 鈴木秀子 (2004)9つの性格 エニアグラムで見つかる「本当の自分」と最良の人間関係PHP文庫, ISBN 978-4-569-66106-3The best known Japanese book on the Enneagram, and the source against which the Japanese wording here was compared.