Semantic Differential Scale

The Semantic Differential Scale is a tool commonly used in linguistics and social psychology to measure social attitudes. Introduced by Osgood, Suci, and Tannenbaum in 1957, it usually employs a seven-point bipolar rating system with opposing adjectives, though some studies use five- or nine-point scales.

The semantic differential technique of Osgood et al. (1957) asks a person to rate a topic on a standard set of bipolar adjectives, meaning pairs with opposite meanings.

Each pair forms its own seven-point scale.

 semantic differential technique
Respondents indicate their perceptions by placing a mark on the continuum between the two opposing words. Each point on the scale represents a more nuanced position between the two extremes.

How it Works

  1. Selection of a Concept: Choose the concept, idea, brand, or item you want to measure. This could be anything from a product, service, brand, to even an abstract idea.

  2. Bipolar Adjective Pairs: Develop a list of bipolar (opposite) adjectives. These pairs will act as the endpoints of your scale, reflecting the dimensions you want to measure (e.g., “Good-Bad”, “Fast-Slow”). A pilot study is typically conducted to determine the most suitable measures for the main study.

  3. Constructing the Scale: Typically, a 5-point, 7-point, or sometimes even a 9-point scale is used. Place your bipolar adjectives at either end of this scale:

    Happy -----------|-----------|-----------|-----------|-----------|-----------|----------- Sad

  4. Survey Deployment: Distribute the scale to your intended audience. Ask participants to place a mark on the continuum between the paired adjectives, indicating their perception or feeling about the concept.

  5. Analysis: After collecting responses, the data can be analyzed by computing average scores or distributions for each bipolar adjective pair. This provides insights into how respondents perceive the concept across different dimensions.

  6. Interpretation: The positions marked by respondents on the scale indicate their attitudes or feelings. For example, if most respondents mark closer to “Happy” for a product, it indicates a positive sentiment towards that product.

Dimensions 

The semantic differential technique reveals three basic dimensions of attitudes: evaluation, potency (i.e., strength), and activity.

Evaluation is concerned with whether a person thinks positively or negatively about the attitude topic (e.g. dirty – clean, and ugly – beautiful).

Potency is concerned with how powerful the topic is for the person (e.g. cruel – kind, and strong – weak).

Activity is concerned with whether the topic is seen as active or passive (e.g. active – passive).

By assessing attitudes across these three dimensions, the Semantic Differential Scale offers a nuanced and multi-faceted view of a respondent’s perception.

For example, a brand might be perceived as strong (Potency) but not necessarily good (Evaluation). This distinction matters for marketers, researchers, and other professionals.

This helps reveal whether a person’s feelings about something match their behavior. For example, someone might love the taste of chocolate (evaluative) yet rarely eat it (activity).

Social psychologists have mostly used the evaluation dimension to measure a person’s attitude because this dimension reflects the affective aspect of an attitude.

The Founding Study: Osgood, Suci and Tannenbaum (1957)

These three dimensions were not simply proposed. Osgood, Suci and Tannenbaum (1957) discovered them, and the study behind them anchors the whole technique.

Aim: to build a general-purpose way of measuring the connotative meaning of any concept, without tying the instrument to one specific attitude object.

Method: large samples of respondents rated many different concepts, including people, objects and abstract nouns, on extensive batteries of seven-point bipolar adjective scales such as good-bad, strong-weak and hot-cold. These ratings were then factor-analysed. Factor analysis groups scales that respondents rate together.

Results: whatever the concept being rated, three factors consistently explained most of the shared variance: evaluation, potency and activity. Evaluation typically explained the largest share and matched the everyday sense of “attitude.”

Conclusion: connotative meaning has a compact, recurring structure.

Because the attitude object is a single word rather than a bespoke statement, the same battery of scales can measure and compare genuinely different concepts.

Examples

Concept Being Evaluated: XYZ Coffee Brand

Please indicate where your perception of XYZ Coffee Brand falls on the following scales:

  1. Taste
    Bland ———–|———–|———–|———–|———–|———–|———– Flavorful

  2. Aroma
    Weak ———–|———–|———–|———–|———–|———–|———– Strong

  3. Price
    Expensive ———–|———–|———–|———–|———–|———–|———– Affordable

  4. Packaging
    Unattractive ———–|———–|———–|———–|———–|———–|———– Attractive

  5. Strength
    Mild ———–|———–|———–|———–|———–|———–|———– Robust

  6. Aftertaste
    Unpleasant ———–|———–|———–|———–|———–|———–|———– Pleasant

Respondents would then mark the continuum between the bipolar adjectives based on their perceptions of the XYZ Coffee Brand. The results will give an insight into how the new coffee brand is perceived across different attributes.

Separating Evaluation from Potency and Activity

A second example shows why three separate dimensions matter. A political candidate might be rated as strong and active without being rated as good.

A voter can see a candidate as powerful and energetic while still disliking them. A single favourable-to-unfavourable question would collapse this distinction and lose it.

The Semantic Differential Scale keeps the three ratings separate. A researcher can then see exactly when perceived strength and dynamism go together with liking, and when they do not.

This has practical value beyond politics. A brand can be perceived as strong and active in advertising tests without customers necessarily liking it. That result tells a marketing team the message lands on potency and activity, but not yet on evaluation, the dimension closest to overall attitude.

Semantic Differential vs the Likert Scale

The Likert scale (Likert, 1932) is the Semantic Differential Scale’s closest rival. It remains more widely used because it is reliable and quick to build.

How the Two Techniques Differ

A Likert scale asks respondents to rate their agreement with a pool of statements about one specific attitude object, on a continuum from strongly agree to strongly disagree.

The Semantic Differential Scale instead uses generic bipolar adjective pairs, such as good-bad or strong-weak, that are not written for any single object.

This generality is the key difference. The same rating form can profile a coffee brand, a political candidate or a policy without being rewritten. A Likert scale’s statements must be rebuilt from scratch for every new object.

Likert items also go through an extra step: item analysis. This keeps only the statements that best separate high-scoring from low-scoring respondents. The Semantic Differential Scale needs no equivalent step, because its adjective pairs are reused unchanged across every concept.

Which Technique to Choose

A Likert scale is the better choice when a researcher wants rich, statement-level detail about why one specific attitude is held.

The Semantic Differential Scale is the better choice when a researcher wants to compare many different attitude objects on one common set of scales. The two techniques are complementary, and many studies use both.

Both techniques also displaced an older method, Thurstone’s method of equal-appearing intervals, which needed a panel of judges to sort statements into eleven categories of favourability.

That step made Thurstone scaling expensive to build, and neither the Likert scale nor the Semantic Differential Scale requires it. The lesson is the same for both.

A researcher studying why customers dislike a product would reach for a Likert scale, since its statements can probe specific reasons. A researcher wanting to rank ten competing products on the same scale would reach for the Semantic Differential Scale instead.

Real-World Applications

Because the Semantic Differential Scale can rate any concept on the same instrument, it has been put to work well beyond the psychology laboratory.

Market and Political Research

Market researchers rate a brand on scales such as modern-traditional or reliable-unreliable, producing a profile that can be tracked before and after an advertising campaign, or compared against a competitor.

The same logic applies to political and social attitude research. One instrument can profile a politician, a policy or a social group on evaluation, potency and activity together.

This captures a perception such as “strong but not well-liked” that a single favourable-unfavourable question would flatten into one score. That gap matters for strategy.

A profile like this can also guide packaging and branding decisions for a specific product line.

A design team might want to shift a product’s perceived potency or activity, making it look bolder or more energetic. That happens without damaging overall evaluation.

Affect Control Theory in Sociology

Osgood’s evaluation-potency-activity dimensions became the measurement backbone of Affect Control Theory (Heise, 2007; MacKinnon & Heise, 2010), a sociological framework built on exactly this technique.

The theory uses the shared connotations attached to social identities, behaviours and settings, called “fundamental sentiments.” This is Osgood’s framework at work. These connotations predict how people interpret social events and which actions feel appropriate within them.

Researchers have built cross-national “cultural dictionaries” of these ratings for thousands of identity and behaviour words. All of them use the semantic-differential method Osgood established.

This body of work continues to grow. Contemporary sociological research on status, identity and social interaction still relies on these Affect Control Theory dictionaries. They predict how an interaction, such as a greeting or a display of anger, is likely to unfold.

Evaluation of the Semantic Differential Scale

The Semantic Differential Scale has real strengths as an attitude-measurement tool, but it also carries a limitation common to every self-report method.

  1. Captures Direction and Intensity: the scale records both which way an attitude points and how strongly it is held, not just a positive-or-negative label.
  2. Social Desirability Bias: as a self-report method, it is vulnerable to respondents giving answers that make them look well-adjusted rather than answers that reflect their genuine attitude.

Captures Direction and Intensity

Most single-item ratings can only say whether an attitude is positive or negative. The Semantic Differential Scale goes further: it records both the direction of an attitude and its intensity. Each bipolar scale offers several rating points rather than a single yes-or-no judgement.

Scoring reflects this too. Analyzing the results typically means calculating a mean score for each bipolar item, turning individual marks into one comparable group-level measure.

This is what makes the scale useful for tracking attitude change over time, or for comparing groups. A shift toward the scale’s midpoint shows weakening intensity even when the overall direction stays the same, a distinction a simple agree-disagree question would miss.

This distinction has real consequences. Two respondents who both feel “negative” about something may hold that feeling with very different force.

Social Desirability Bias

An attitude scale is meant to give a valid, accurate measure of a person’s social attitude.

But anyone who has ever “faked” an attitude scale knows self-report measures have shortcomings. The most common problem is social desirability.

This is the tendency to give answers that appear well-adjusted, unprejudiced, open-minded and democratic rather than answers that reflect a genuine attitude. This affects validity most heavily on scales measuring attitudes toward race, religion or sex.

Someone who privately holds a negative attitude toward a group may not want to admit that, even to themselves. Responses on attitude scales are therefore not always 100% valid.

Researchers working on sensitive topics increasingly pair the Semantic Differential Scale with implicit alternatives, such as evaluative priming or the Implicit Association Test, rather than relying on self-report alone.

Contemporary Research

Research since 2015 updates this technique’s evidence base along two threads. One asks whether ratings are comparable across demographic groups; the other asks whether computational models can reproduce them.

Chapman, Gardner and Lyons (2022) recruited 94 young adults.

They were 47 men and 47 women, aged 18-39, and rated 120 words on the three original dimensions.

Aim: to test whether men and women, using the identical rating instrument, differ systematically in the connotative meaning they assign to emotionally loaded words.

Method: participants rated positive-, negative- and neutral-valence words on all three dimensions. Ratings were then compared by gender within each valence category.

Results: young women rated negative-valence words consistently more negatively than young men did, across all three dimensions. No comparable difference emerged for positive- or neutral-valence words.

Conclusion: gender moderates connotative processing for negative content. This matters when comparing ratings across mixed-gender samples.

A newer strand asks whether large language models can reproduce human ratings on these dimensions. Combs and colleagues (2025) compared commercial language models’ outputs against human benchmark data.

The comparison was not meant to settle anything. It was useful mainly for flagging cultural-bias risks. Nearly seventy years after 1957, the three-dimensional structure is still the benchmark new tools are checked against.

Key Takeaways

  • What it Measures: the Semantic Differential Scale rates the connotative meaning of a concept — how it feels — using pairs of opposite adjectives on a seven-point scale.
  • EPA Dimensions: ratings reduce to three recurring factors: Evaluation (good-bad), Potency (strong-weak) and Activity (active-passive).
  • Generic Scales: because the same adjective pairs work for any concept, the identical instrument can rate and compare completely different things, unlike the object-specific Likert scale.
  • Self-Report Limits: as with any self-report measure, ratings can be distorted by social desirability, especially on sensitive topics.
  • Cross-Cultural Stability: the three-dimensional structure has been found across roughly two dozen language communities, suggesting it reflects something general about how humans organise affective meaning.

References

  • Al-Hindawe, J. (1996). Considerations when constructing a semantic differential scale.
  • Chapman, R. M., Gardner, M. N., & Lyons, M. (2022). Gender differences in emotional connotative meaning of words measured by Osgood’s semantic differential techniques in young adults. Humanities and Social Sciences Communications, 9, Article 116. https://doi.org/10.1057/s41599-022-01126-3
  • Combs, A., Dametto, D., Blaison, C., Schröder, T., Hoey, J., & Smith-Lovin, L. (2025). Affective connotations according to LLMs: Implications for meaning measurement and cultural bias. Cognition and Emotion. https://doi.org/10.1080/02699931.2025.2568551
  • Heise, D. R. (1969). Some methodological issues in semantic differential research. Psychological Bulletin72(6), 406.
  • Heise, D. R. (1970). The semantic differential and attitude researchAttitude measurement4, 235-253.
  • Heise, D. R. (2007). Expressive order: Confirming sentiments in social actions. Springer.
  • Garland, R. (1990). A comparison of three forms of the semantic differentialMarketing Bulletin1(1), 19-24.
  • Likert, R. (1932). A technique for the measurement of attitudes. Archives of Psychology, 22(140), 5-55.
  • MacKinnon, N. J., & Heise, D. R. (2010). Self, identity, and social institutions. Palgrave Macmillan.
  • Osgood, C. E., Suci, G. J., & Tannenbaum, P. H. (1957). The measurement of meaning. University of Illinois Press.
 

Olivia Guy-Evans, MSc

BSc (Hons) Psychology, MSc Psychology of Education

Associate Editor for Simply Psychology

Olivia Guy-Evans is a writer and associate editor for Simply Psychology, where she contributes accessible content on psychological topics. She is also an autistic PhD student at the University of Birmingham, researching autistic camouflaging in higher education.


Saul McLeod, PhD

Chartered Psychologist (CPsychol)

BSc (Hons) Psychology, MRes, PhD, University of Manchester

Saul McLeod, PhD, is a qualified psychology teacher with over 18 years of experience in further and higher education. He has been published in peer-reviewed journals, including the Journal of Clinical Psychology.