Minnesota Multiphasic Personality Inventory (MMPI)

The term MMPI in personality assessment refers to the Minnesota Multiphasic Personality Inventory. This standardized psychometric test identifies personal, social, and behavioral issues in psychiatric patients.

Key Takeaways

  • What It Measures: The MMPI is a psychological instrument that assesses personality traits and forms of psychopathology.
  • Who Administers It: A trained psychologist or mental health professional gives the test and uses ten clinical scales to help diagnose mental health or other clinical issues.
  • Origins: Starke Hathaway and J.C. McKinley first developed the MMPI in 1939; it has since been revised several times, including a version for adolescents.
  • Versions: The original MMPI, developed in the 1940s, is still in use and has 550 true/false items. The MMPI-2, introduced in 1989, has around 567 items.
  • Scale Types: Clinicians mainly use the clinical and validity scales, plus a content scale and several supplemental scales that add further insight.
  • Use and Criticism: The MMPI-2 is used from treatment planning and hiring to law enforcement and marital counseling. Scholars have raised concerns about racial disparities in scoring and about the scientific validity of some scales.
Mind behavior, mental mindset types or models Concept With icons. Cartoon Vector People Illustration
The Minnesota Multiphasic Personality Inventory (MMPI) works by assessing an individual’s psychological traits and providing insights into various mental health conditions by measuring their responses to a set of standardized questions or statements.

History of the MMPI

The Minnesota Multiphasic Personality Inventory (MMPI) is probably the most widely used multidimensional tool to help diagnose mental health disorders. This test was first developed in 1939 by clinical psychologist Starke Hathaway and neuropsychiatrist J.C. McKinley at the University of Minnesota, and it was first published in 1943.

This tool is unique. It was created to measure psychopathology specifically, not to assess personality in general.

To develop it, the team constructed clinical scales endorsed by patients who had been diagnosed with certain mental disorders (Hathaway & McKinley, 1940).

Unlike other personality tests, the MMPI was not based on any particular prevailing theories about personality, such as the five-factor model or 16 personalities. It simply measures where someone falls on 10 clinical scales. Clinicians use those scores to diagnose the patient and select the proper treatment they need.

The MMPI developed from hundreds of true/false questions that were felt would be useful in identifying personality dimensions. These questions were then given to people suffering from a variety of psychological disorders and to a group that did not suffer from any disorder.

They then compared the two groups’ answers. Questions that most people with a particular illness answered differently from those identified as normal were kept.

Consequently, if a person answered many questions the way paranoid people did, it was highly likely that the individual was also paranoid.

How the Test Has Changed

Although this atheoretical approach allowed the MMPI to meaningfully capture traits that were directly related to psychopathology, the initial version was widely criticized.

Specifically, the original control group had a very small sample size, and it was primarily young, white, and married people from the rural Midwest.

Additionally, the MMPI was attacked for using incorrect terminology and not including a wide enough range of mental health issues, such as suicidal tendencies and drug abuse (Gregory, 2004).
As a result, the MMPI went through multiple stages of revision.

MMPI-2

In 1989, the first major revision of this tool occurred, producing the MMPI-2. The initial inventory had not been tested on a representative sample, so the MMPI-2 was standardized on 2,600 individuals from more diverse backgrounds (Gregory, 2004).

Additionally, some of the items on the test were revised and some sub-scales were introduced to better help the clinicians administering the MMPI interpret the results.

The MMPI-2 contains 567 true/false items. It is meant for adults and typically takes around 60 to 90 minutes to complete.

It is designed for all adults over the age of 18 and requires a sixth-grade reading level (Gregory, 2004).

Although this specific test isn’t made for those who are younger than 18, an adolescent version of the MMPI has also been created.

MMPI-A

A couple of years after the MMPI-2 was released, the MMPI-A (A for adolescents) was developed to include individuals aged 14 to 18.

After empirical testing revealed that 12 and 13-year-olds could not sufficiently understand the questions, 14 was deemed the cutoff for this personality inventory.

To this day, there is still no version for children younger than 14. The MMPI-A contains 478 items, and adolescents are scored along the same 10 scales that are used for the MMPI-2 (Kaemmer, 1992).

But why does there need to be a separate MMPI for adolescents? Doesn’t the one for adults suffice?

The MMPI-A closely resembles the adult version of the test, but clinicians were concerned that the MMPI-2’s content was inadequate for adolescents. The items did not cover topics relevant enough to teenagers, such as relations with peers or school.

They also feared that there was a lack of appropriate norms which are necessary for addressing how an individual might be deviating from such norms.

Adolescence is a critical period of development. The norms for how a 14-year-old is expected to act would not be the same as for an adult whose brain is finished maturing.

When trying to decide whether to use adult norms that highly pathologized children or recently-published adolescent norms which hadn’t been extensively reviewed, psychologists could not even reach a consensus.

A separate adolescent version therefore made sense (Kaemmer, 1992).

To test the new MMPI-A and establish new norms, psychologists recruited 1,620 individuals for the normative sample and 713 individuals for the clinical sample.

This new version was praised for including more relevant item content, being shorter, and having a high degree of validity (Kaemmer, 1992).

Like the MMPI-2, the MMPI-A continues to be criticized.

One of the major criticisms is that the clinical sample used to establish the norms was not a representative sample. All of the individuals were in treatment facilities in Minneapolis, Minnesota, so it did not reflect geographic areas around the globe.

Scholars also argue the 10 clinical scales overlap too much. And even though the MMPI-A is shorter than the original MMPI, they say it is still too long and advanced for adolescents (Whitcomb & Merrell, 2013). As a result, a new scale was established.

MMPI-2 RF

Ideas and theories are constantly evolving, and there are always ways to improve upon methodological flaws in the research design.

The research world is an endless cycle of testing, publishing, receiving critiques, and retesting. And the development of the MMPI followed this same pattern. After both the MMPI-2 and the MMPI-A received backlash from the academic community, both received revisions.

The MMPI-2 RF (restructured form) was published in 2008, inspired by a restructuring of the ten clinical scales that was done in 2003 (Tellegen et al., 2003).

This new version has been extensively tested in the empirical setting and is able to differentiate between different clinical symptoms and broader diagnoses.

It only has 338 questions, significantly fewer than the MMPI-2, and takes around 35-50 minutes to complete.

MMPI-A RF

The MMPI-A RF was first published in 2016 and sought to address the many criticisms that the original adolescent multiphasic inventory received.

Namely, it only contains 241 true-false items – less than half the number of items of the original MMPI-A to help combat the challenges of adolescent attention span and concentration.

Additionally, this version tries to do a better job of increasing discriminant validity, which occurs when non-overlapping factors do, in fact, not overlap (Handel, 2016).

The MMPI-A RF is one of the most commonly used psychological tools among the adolescent population (Whitcomb & Merrell, 2013).

MMPI-3

The most recent revision, the MMPI-3, was published in 2020.

It updated the norms again on a new, nationally representative U.S. sample, refreshed and expanded the item content, and reorganized the scale structure inherited from the RF.

A key practical question with every revision is comparability: can data collected on an older form be used to estimate scores on the newer one?

A 2024 study developed and validated a proration method for deriving MMPI-3 scores from MMPI-2 and MMPI-2-RF item responses (Brown, Menton, & Ben-Porath, 2024).

The method fully scored 16 of the MMPI-3’s 52 scales and reliably prorated a further 24, for 40 of the 52 scales in total.

This lets archival MMPI-2 and MMPI-2-RF datasets remain useful for studying most MMPI-3 scales without recollecting data.

Administration

The MMPI-2 restandardization includes 567 items and takes approximately 60 to 90 minutes to complete. Like the MMPI, it has 10 clinical scales, each with roughly 32 to 78 items.

Who Administers the MMPI and Why

The vast majority of people obtain a few scores on each scale. Those administering the test are looking for a higher-than-average score on a scale to indicate a potential problem.

The MMPI is copyrighted by the University of Minnesota, so clinicians must pay to administer and use it.

There are several reasons for administering the MMPI. A clinician will typically give it to help develop treatment plans for a patient and assist with differential diagnosis.

It can also be used as part of the therapeutic assessment procedure or to help answer legal questions (Butcher & Williams, 2009). Another common use is to screen job candidates during personnel selection, especially in careers such as law enforcement.

The MMPI may also be used during college, career, and marital counseling, as well as in child custody disputes and substance abuse programs (Ben-Porath & Tellegen, 2008).

The MMPI can only be administered and interpreted by psychologists who have been extensively trained in it. It can be done online or with a physical booklet.

After a participant takes the test, an evaluator prepares an interpretative report based on their responses.

Scoring: T-Scores and Codetypes

Raw scores are converted into normalized T-scores, ranging from 30 to 120, with a population average of 50. On the modern forms, a T-score of around 65 and above is treated as clinically significant (Framingham, 2016).

Over time, a set of standard clinical profiles, or “codetypes,” have emerged. A codetype is when two clinical scales have high T-scores.

For example, a 2-7 codetype means a participant scored high on both scale 2 (depression) and scale 7 (psychasthenia or OCD), signaling both depression and anxiety. Many common codetypes like this one have been identified and are well understood by clinicians.

The MMPI-2 is never administered in a vacuum: a participant’s background is taken into account, and it is typically not the only evaluative tool a person is given. Even so, it remains a valuable tool for identifying psychopathology.

10 Clinical Scales

The MMPI is designed to evaluate the thoughts, feelings, attitudes, and behaviors that make up an individual’s personality.

The test is administered by a trained clinician, typically a psychologist or psychiatrist, who relies on a series of clinical scales to interpret the results (Framingham, 2016).

These clinical scales help illustrate certain forms of psychopathology that an individual may have. The 10 clinical scales are as follows:

  1. Hypochondriasis (Hs): excessive concern about bodily complaints, typically focused on the back and abdomen, that persist even when medical tests are negative. A high score suggests health worry is interfering with daily life and relationships. 32 items.
  2. Depression (D): clinical depression marked by low morale, hopelessness, loss of interest in activities, worthlessness, and poor concentration. 57 items.
  3. Hysteria (Hy): reactions to stress shown through physical complaints and denial of problems, examined across poor physical health, shyness, cynicism, and headaches. 60 items.
  4. Psychopathic Deviate (Pd): social maladjustment and rule-breaking, including conflict with family and authority figures, self- and social alienation, and boredom. A high score may indicate antisocial tendencies. 50 items.
  5. Masculinity/Femininity (Mf): interests, hobbies, and aesthetic preferences historically keyed to how closely a person conforms to stereotyped masculine or feminine roles. 56 items.
  6. Paranoia (Pa): interpersonal sensitivity, moral self-righteousness, and suspiciousness, including delusional thoughts, grandiose thinking, and feelings of persecution at the extreme. 40 items.
  7. Psychasthenia (Pt): an outdated term for what is now called obsessive-compulsive disorder (OCD), covering compulsive behaviors, abnormal fears, self-criticism, and anxiety. 48 items.
  8. Schizophrenia (Sc): unusual thinking and perception, social alienation, and difficulty concentrating; it has the greatest number of items of any scale. It indexes deviant experience rather than diagnosing the disorder outright. 78 items.
  9. Hypomania (Ma): elevated, unstable mood, psychomotor overactivity, impulsivity, rapid speech, and irritability. 46 items.
  10. Social Introversion (Si): tendency toward social withdrawal versus outgoing sociability, including competitiveness, compliance, and timidity. Added later than the original nine scales. 69 items.

MMPI Clinical Scales

Together, these 10 dimensions comprise the clinical scales of the MMPI. Although there are many supplementary scales, these are the main 10 that guide a clinician’s evaluation of an individual. However, to accompany these 10 scales, a series of validity scales are also used to ensure the accuracy of the results.

Validity Scales

Validity refers to how accurately a method is actually measuring what it intends to measure. A common analogy for validity involves a dart board.

If all of the arrows are close to the bullseye, that is an example of strong validity because all of the arrows are close to what they are supposed to be close to. Even if some are slightly too high and some are slightly too low, if all of the arrows are close to the middle, that indicates high validity.

For the MMPI, there are not only studies that investigate the validity of the MMPI as a tool, but the instrument itself has its own validity scales to ensure that the test taker’s answers aren’t over-reported, under-reported, or flat out dishonest. In other words, the following scales help ensure that the respondent’s answers reflect what they should actually be reflecting (Framingham, 2016):

  1. Lie (L): flags attempts to look unrealistically good by denying common, minor faults nearly everyone admits to. This scale contains 15 items.
  2. F: flags rare, unusual, or symptom-heavy answers, which may reflect genuine severe disturbance, random responding, or an attempt to “fake bad.” This scale contains 60 items.
  3. Back F (Fb): the same idea as the F scale, but scored from items in the second half of the test, to catch fatigue or a change in response pattern partway through. This scale has 40 items.
  4. K: a subtler defensiveness index than L, measuring self-control and guarded self-presentation; in some scales it statistically corrects clinical scores upward. This scale is composed of 30 items.
  5. Cannot Say (?): counts how many items a person leaves unanswered or marks both true and false. More than 30 unanswered items may invalidate the profile.
  6. TRIN and VRIN: catch careless or contradictory answering. TRIN (True Response Inconsistency) flags an indiscriminate tendency to answer all true or all false; VRIN (Variable Response Inconsistency) flags random, inconsistent answers to similar item pairs.
  7. Fp: separates genuine severe pathology from exaggeration by using items rare even among psychiatric patients, not just the general population. This scale has 27 items.
  8. FBS: also called the Fake Bad Scale or Symptom Validity Scale; designed to detect intentional over-reporting of symptoms, mainly in personal-injury and disability cases. This scale has 43 items.
  9. S: the Superlative Self-Presentation scale, used mainly in non-clinical settings like employment screening, measuring positive self-presentation through 50 questions about serenity, morality, and patience.

MMPI Validity Scales

Empirical Studies Validating MMPI

As mentioned, high validity occurs when a tool measures what it claims to be measuring. So, in the case of the MMPI, high validity means that this inventory is actually measuring different types of psychopathology as it claims to be.

Take someone diagnosed with depression through other means. If the MMPI also revealed depressive traits in that individual, it would be deemed valid.

Conveniently, validity can be tested empirically. Several research studies have sought to examine the degree to which the MMPI is actually a valid tool.

The most-cited evidence on MMPI validity is a 1999 meta-analysis comparing it with the Rorschach inkblot test (Hiller et al., 1999).

Aim: to estimate and compare the criterion validity of the MMPI and the Rorschach across the published literature. This tested a common assumption. Is the objective MMPI necessarily more valid than the projective Rorschach?

Method: the researchers synthesised validity coefficients from many studies correlating each instrument’s scores with outcomes. A coefficient of 0.21 to 0.35 is labeled useful.

Results: the two instruments produced comparable average validity coefficients. The MMPI averaged 0.3 and the Rorschach 0.29. The MMPI had larger coefficients for studies using psychiatric diagnoses and self-report measures, while the Rorschach scored higher on studies using objective criterion variables.

Conclusion: the MMPI is a valid instrument, but its validity is moderate and roughly on par with a well-scored Rorschach. A coefficient near 0.3 means the test carries real, useful information while leaving much unexplained. So a profile is one input, not a verdict.

Another study measured the extent to which the MMPI can discriminate between sex offenders and controls.

In the study, 479 men took the MMPI and were scored on scales measuring sexual behavior and deviance, substance abuse, personality, violence, and more.

The MMPI screened sex offenders better than chance. This supports its validity as a tool for correctly identifying sex offenders (Langevin et al., 1990).

Some studies even test the validity scales themselves. For example, one study looked at the validity of the FBS scale which is used to identify individuals who are over-reporting somatic symptoms.

The researchers measured associations between scores on the MMPI-FBS scale and a symptom validity test (SVT) in 127 criminal defendants and 141 civil claimants.

The researchers found that scores on the FBS were, in fact, associated with SVT failure (incorrectly reporting symptoms) in both the criminal and civil defendants (Wygant et al., 2007).

This supports the FBS’s utility. It flags inaccurate reporting of both somatic and cognitive complaints.

Another study observed the validity of the F and Fp scales in detecting psychopathology in a criminal forensic setting.

One hundred twenty-five criminal defendants were given a Structured Interview of Reported Symptoms (SIRS) and the MMPI-2 RF. The two MMPI-2 RF over-reporting scales properly differentiated between defendants who were genuinely ill and those who were faking it, supporting the scales’ validity (Sellbom et al., 2010).

Together, these studies show the MMPI-2 is valid. Not only does it have validity scales in place, but these scales themselves are valid too.

Having this kind of confirmation allows this tool to be viewed as highly credible and effective by the research world.

Empirical Research

Research on the MMPI-2 extends beyond validity testing. Several studies have examined the tool across different settings and populations, from clinical samples to college students to workplaces.

Clinical and Occupational Populations

One study relied on the MMPI-2 to identify key disorders among battered women in transition. Of the 31 women evaluated, 90% scored high on the psychopathy, paranoia, and schizophrenia scales (Khan et al., 1993).

The MMPI-2 thus guides both diagnosis and intervention.

A separate study examined the workplace, where bullying, defined as exposure to psychological violence and harassment, places an unnecessary strain on employees and their productivity. Researchers used the MMPI-2 to identify psychological correlates of bullying among 85 former and current victims.

Bullied victims showed an elevated MMPI-2 profile, and the profile’s shape tracked the type and intensity of bullying they experienced (Matthiesen & Einarsen, 2001).

A further study looked at a different population: college students. Researchers administered the MMPI-2 to 515 male and 797 female college students across four universities.

The group responded to the test much like the normative sample, supporting the idea that the MMPI-2’s norms generalize to college populations (Butcher et al., 1990).

These are just a couple of examples from an expansive research literature: researchers have targeted varying populations, settings, and even different versions of the test.

Critical Evaluation

Although the MMPI-2 is a widely used and highly valid test, it does not come without its criticisms. Here are the three main concerns raised about it:

  1. Cultural and Demographic Bias: non-white respondents score higher on average across many scales, a pattern traced to the test’s original, mostly white and Minnesotan reference groups, not to real differences in psychopathology (McCreary & Padilla, 1977).
  2. FBS Validity Controversy: Butcher and colleagues argue the Fake Bad Scale doesn’t cleanly measure over-reporting and can be biased against women, people with disabilities or trauma histories, and psychiatric inpatients (Butcher et al., 2008).
  3. RF vs. MMPI-2 Debate: Butcher and Williams argue the shorter MMPI-2-RF isn’t a safe substitute for the MMPI-2 in high-stakes evaluations like child custody, though critics call their examples cherry-picked (Butcher & Williams, 2012).

Cultural and Demographic Bias

Studies of incarcerated people, medical patients, and high school students have repeatedly found that Black and other non-white respondents score higher than white respondents across many clinical scales. The gap runs several T-score points on average (McCreary & Padilla, 1977).

Applied uncritically at a fixed cut-score, that pattern risks over-diagnosing minority respondents with psychopathology they don’t actually have.

The likely explanation lies in how the test was built. Empirical keying bakes the response tendencies of the original criterion and normative groups directly into the scales.

Those groups were drawn almost entirely from a white, rural Minnesotan population of the 1930s and 1940s.

Later re-norming, including the MMPI-2’s larger and more diverse standardization sample, has reduced the disparity but not eliminated it. The concern is most consequential precisely where the MMPI is common: high-stakes forensic and employment evaluations.

FBS Validity Controversy

Butcher, Gass, Cumella, Kally, and Williams (2008) criticize the FBS, or Fake Bad Scale, on two separate grounds: psychometric and fairness.

It was later renamed the Symptom Validity Scale. Psychometrically, they argue it does not cleanly isolate over-reporting from other causes of an elevated score.

They also say the criteria used to define malingering, exaggerating or feigning illness, are not consistently replicable across studies.

On fairness, they argue the scale can be biased against women, trauma survivors, and people with genuine disabilities or medical illness.

Real symptoms may inflate the scale just as faking would.

Because the FBS is widely used to flag symptom exaggeration in personal-injury litigation, a validity scale that mislabels genuine sufferers as fakers is not a minor technicality.

It can affect who receives compensation, treatment, or credibility in a legal proceeding.

RF vs. MMPI-2 Debate

In a separate paper, Butcher and Williams (2012) argue the shorter MMPI-2-RF is not a viable replacement for the MMPI-2 in high-stakes evaluations such as child custody.

The restructured scales, they say, lose information that the longer instrument’s decades of interpretive research still carry. Many forensic and custody evaluators continue to prefer the MMPI-2 over the RF for exactly this reason.

However, the scientific community has pushed back hard against Butcher and Williams’ specific critique, calling it misleading and its examples cherry-picked rather than representative.

The debate illustrates a broader tension in the MMPI’s history. Each revision gains construct clarity and efficiency, but has to earn its place against an older form’s accumulated evidentiary weight before clinicians and courts fully adopt it.

Contemporary Research

The strongest recent evidence on the MMPI concerns its signature capability: detecting when a respondent is distorting their answers. This might mean exaggerating symptoms (over-reporting, “faking bad”) or minimizing them (under-reporting, “faking good”).

This matters most in high-stakes settings where a person has an incentive to distort, including personal-injury litigation and pre-employment screening.

Aim: a meta-analysis by Ingram and Ternes (2016) set out to quantify how well the MMPI-2-RF’s five over-reporting validity scales distinguish honest responders from people feigning or exaggerating symptoms. It also tested what moderates that performance.

Method: the researchers located 25 experimental and quasi-experimental studies comparing validity-scale scores between honest and over-reporting groups, then ran moderated meta-analyses testing several moderators.

Results: every over-reporting scale was an effective general discriminator, with mean effect sizes ranging from about 1.08 (Symptom Validity, FBS-r) to about 1.43 (Infrequent Psychopathology, Fp-r). Fp-r and the response-bias scale were the least affected by moderating factors, and the honest-versus-feigning separation was narrower than traditional cut-scores imply.

Conclusion: the RF validity scales reliably detect patterned over-reporting. Their thresholds should be applied with attention to base rates and the evaluation context rather than mechanically (Ingram & Ternes, 2016).

Known-groups studies in criminal-forensic settings reach the same conclusion using real evaluees, not instructed simulators (Sellbom et al., 2010).

Studies linking MMPI-2 validity scores to independent symptom-validity tests agree. The same numerical elevation carries different weight in a treatment-seeking clinic than in a compensation-seeking forensic referral (Wygant et al., 2007).

Together, these converging designs make response-validity assessment the MMPI’s best-replicated contemporary contribution: not diagnosis itself, but telling the clinician how much to trust the rest of the profile.

In general, despite these open questions about bias and interpretation, the MMPI-2 remains widely accepted by the scientific community.

FAQs

How many different scales are used to compile a profile from the Minnesota multiphasic personality inventory (MMPI)?

The Minnesota Multiphasic Personality Inventory (MMPI) uses 10 clinical scales to compile a profile. However, in its latest version, MMPI-2-RF, there are 50 scales, including validity, substantive, and supplementary, to provide more detailed personality and psychopathology assessments.

What is one main difference between the Minnesota multiphasic personality inventory (MMPI) and the Rorschach inkblot test?

One main difference is that the Minnesota Multiphasic Personality Inventory (MMPI) is a self-report inventory with specific, structured questions, while the Rorschach Inkblot Test is a projective test where individuals interpret ambiguous inkblots, revealing unconscious desires and conflicts.

What is the purpose of the MMPI?

The purpose of the Minnesota Multiphasic Personality Inventory (MMPI) is to assess and measure various aspects of an individual’s personality, mental health, and psychopathology, providing valuable insights for clinical diagnosis, treatment planning, and research in fields such as psychology, psychiatry, and counseling.

References

Ben-Porath, Y. S. (2012). Interpreting the MMPI-2-RF. Minneapolis: University of Minnesota Press.

Ben-Porath, Y. S., & Tellegen, A. (2008/2011). MMPI-2-RF (Minnesota Multiphasic Personality Inventory-2 Restructured Form) manual for administration, scoring, and interpretation. Minneapolis: University of Minnesota Press.

Ben-Porath, Y. S., & Tellegen, A. (2008). Empirical correlates of the MMPI–2 Restructured Clinical (RC) scales in mental health, forensic, and nonclinical settings: An Introduction. Journal of Personality Assessment, 90(2), 119-121.

Brown, J. R., Menton, W. H., & Ben-Porath, Y. S. (2024). Development and validation of a method for deriving MMPI-3 scores from MMPI-2/MMPI-2-RF item responses. Psychological Assessment, 36(11), 665-679. https://doi.org/10.1037/pas0001341

Butcher, J. N., Gass, C. S., Cumella, E., Kally, Z., & Williams, C. L. (2008). Potential for bias in MMPI-2 assessments using the Fake Bad Scale (FBS). Psychological Injury and Law, 1(3), 191-209.

Butcher, J. N., Graham, J. R., Dahlstrom, W. G., & Bowman, E. (1990). The MMPI-2 with college students. Journal of Personality Assessment, 54(1-2), 1-15.

Butcher, J. N., & Williams, C. L. (2009). Personality assessment with the MMPI‐2: Historical roots, international adaptations, and current challenges. Applied Psychology: Health and Well‐Being, 1(1), 105-135.

Butcher, J. N., & Williams, C. L. (2012). Problems with using the MMPI–2–RF in forensic evaluations: A clarification to Ellis. Journal of Child Custody, 9(4), 217-222.

Framingham, J. (2016). Minnesota Multiphasic personality Inventory (MMPI). Retrieved from https://psychcentral.com/lib/minnesota-multiphasic-personality-inventory-mmpi#What-Does-the-MMPI-2-Test?

Gregory, R. J. (2004). Psychological testing: History, principles, and applications. Allyn & Bacon.

Handel, R. W. (2016). An introduction to the Minnesota multiphasic personality inventory-adolescent-restructured form (MMPI-A-RF). Journal of clinical psychology in medical settings, 23(4), 361-373.

Hathaway, S. R., & McKinley, J. C. (1940). A multiphasic personality schedule (Minnesota): I. Construction of the schedule. The Journal of Psychology, 10(2), 249-254.

Hathaway, S. R., & McKinley J. C. (1942). Manual for the Minnesota Multiphasic Personality Inventory. Minneapolis: University of Minnesota Press.

Hathaway, S. R., & McKinley, J. C. (1943). The MMPI. Minneapolis.

Hiller, J. B., Rosenthal, R., Bornstein, R. F., Berry, D. T., & Brunell-Neuleib, S. (1999). A comparative meta-analysis of Rorschach and MMPI validity. Psychological Assessment, 11(3), 278.

Ingram, P. B., & Ternes, M. S. (2016). The detection of content-based invalid responding: A meta-analysis of the MMPI-2-Restructured Form’s (MMPI-2-RF) over-reporting validity scales. The Clinical Neuropsychologist, 30(4), 473-496. https://doi.org/10.1080/13854046.2016.1187769

Kaemmer, B. (1992). Minnesota Multiphasic Personality Inventory-Adolescent Version (MMPI-A): Manual for administration, scoring and interpretation.

Khan, F. I., Welch, T. L., & Zillmer, E. A. (1993). MMPI-2 profiles of battered women in transition. Journal of Personality Assessment, 60(1), 100-111.

Langevin, R., Wright, P., & Handy, L. (1990). Use of the MMPI and its derived scales with sex offenders: I. Reliability and validity studies. Annals of Sex Research, 3(3), 245-291.

Matthiesen, S. B., & Einarsen, S. (2001). MMPI-2 configurations among victims of bullying at work. European Journal of work and organizational Psychology, 10(4), 467-484.

Sellbom, M., Toomey, J. A., Wygant, D. B., Kucharski, L. T., & Duncan, S. (2010). Utility of the MMPI–2-RF (Restructured Form) validity scales in detecting malingering in a criminal forensic setting: A known-groups design. Psychological Assessment, 22(1), 22.

McCreary, C., & Padilla, E. (1977). MMPI differences among black, Mexican‐American, and white male offenders. Journal of Clinical Psychology, 33(S1), 171-177.

Tellegen, A., Ben-Porath, Y. S., McNulty, J. L., Arbisi, P. A., Graham, J. R., & Kaemmer, B. (2003). MMPI-2 Restructured Clinical (RC) scales: Development, validation, and interpretation.

Tellegen, A., & Ben-Porath, Y. S. (2008/2011). MMPI-2-RF (Minnesota Multiphasic Personality Inventory-2 Restructured Form) technical manual. Minneapolis. University of Minnesota Press.

Tupes, E. C., & Christal, R. E. (1992). Recurrent personality factors based on trait ratings. Journal of personality, 60(2), 225-251.

Whitcomb, S., & Merrell, K. W. (2013). Behavioral, social, and emotional assessment of children and adolescents. Routledge.

Wygant, D. B., Sellbom, M., Ben-Porath, Y. S., Stafford, K. P., Freeman, D. B., & Heilbronner, R. L. (2007). The relation between symptom validity testing and MMPI-2 scores as a function of forensic evaluation context. Archives of Clinical Neuropsychology, 22(4), 489-499.

Further Reading

Saul McLeod, PhD

BSc (Hons) Psychology, MRes, PhD, University of Manchester

Chartered Psychologist (CPsychol)

Saul McLeod, PhD, is a qualified psychology teacher with over 18 years of experience in further and higher education. He has been published in peer-reviewed journals, including the Journal of Clinical Psychology.


Olivia Guy-Evans, MSc

BSc (Hons) Psychology, MSc Psychology of Education

Associate Editor for Simply Psychology

Olivia Guy-Evans is a writer and associate editor for Simply Psychology, where she contributes accessible content on psychological topics. She is also an autistic PhD student at the University of Birmingham, researching autistic camouflaging in higher education.

Charlotte Ruhl

Harvard Psychology Graduate

BA (Hons) Psychology, Harvard University

Charlotte Ruhl graduated from Harvard University with a degree in Psychology and African American Studies. During her studies she worked at Harvard's Implicit Social Cognition Lab under Dr. Mahzarin Banaji, researching implicit racial bias and outgroup exposure, and contributed to the Decision Science Lab administering studies in behavioural economics and social psychology.