Conversation Analysis (CA) focuses on how language is used in interaction, rather than simply what is being said. CA researchers recognize that conversation is orderly and that this orderliness can be observed and analyzed.
One of the goals of CA is to describe the procedures that people use to produce and understand conversation.
CA aims to understand how people use language to communicate and relies heavily on the analysis of naturally occurring conversations.
It moves beyond simply interpreting words. CA also considers nonverbal cues, turn-taking patterns, and how participants’ actions shape and are shaped by their social roles.
Through detailed examination of real-world conversations, conversation analysis illuminates how individuals use language to construct meaning, exercise power, and navigate the intricacies of social interactions in various settings.
Who introduced conversation analysis?
Conversation analysis was developed by the sociologist Harvey Sacks and his close associates Emanuel Schegloff and Gail Jefferson. They worked at UCLA. This was in the 1960s and early 1970s.
Sacks, Schegloff, and Jefferson laid the groundwork for conversation analysis through their pioneering work on the structure and organization of everyday talk.
They studied recordings of naturally occurring conversations, focusing on the sequential organization of talk and turn-taking. This includes repair mechanisms (how speakers deal with problems of speaking, hearing, or understanding) and the social actions performed through talk.
Some of their seminal works include:
- Sacks, H., Schegloff, E. A., & Jefferson, G. (1974). A simplest systematics for the organization of turn-taking for conversation. Language, 50(4), 696-735.
- Schegloff, E. A., Jefferson, G., & Sacks, H. (1977). The preference for self-correction in the organization of repair in conversation. Language, 53(2), 361-382.
These studies, among others, established conversation analysis as a distinct field within the social sciences, providing a method for analyzing the intricate details of social interaction through talk.
When to use conversation analysis
Conversation Analysis (CA) helps researchers understand the meanings of real language and how people speak in natural settings, going beyond just the words themselves. CA is particularly useful in analyzing:
- Social interaction: CA examines how people use spoken language to interact, focusing on aspects like turn-taking, interruptions, pauses, intonation, and word stress.
- This analysis helps reveal the procedures and structures underlying everyday conversations.
- Power Dynamics: The ability to control who speaks, and when, can reveal power dynamics in conversations.
- For instance, in a courtroom setting, the attorney controls the conversation by asking sequences of questions the witness is expected to answer.
- Institutional Talk: CA observes how participants use talk to navigate formal institutional settings, such as courts or interviews.
- These settings have defined roles, and the structure of turns is often pre-determined. In less formal institutional settings, such as medical or business environments, participants have more flexibility in their roles.
- However, conversation analysis reveals asymmetries in these interactions. For example, in doctor-patient interactions, while the conversation might appear conversational at times, the physician typically controls the topics and determines the outcome of the discussion.
Researchers often choose CA because it relies on naturally occurring data, such as recordings of conversations. CA considers the context of the conversation crucial for understanding the meaning of what is said.
CA vs discourse analysis
CA is distinct from discourse analysis (DA). While both explore language in use, they have different focuses:
- Discourse analysis examines a broader range of language use, considering the meaning of words, intentions, and underlying assumptions within a wider social and cultural context.
- Conversation Analysis zooms in on the non-verbal aspects of speech, such as pauses, intonation, and emphasis, to uncover the subtle ways meaning is created in interactions.
The choice between using CA or DA, or employing them together, depends on the research question.
For example, a researcher might use CA to study how doctors and patients negotiate treatment decisions or how lawyers use language to influence a jury.
Sequential organization of talk
Sequencing refers to the way in which turns and actions in a conversation are ordered and related to each other. Let’s look at how each of these elements relates to sequencing:
- Turn-taking: This refers to the way speakers alternate in taking turns in a conversation. The sequence of turns is a fundamental aspect of the organization of conversation.
- Adjacency pairs: These are pairs of utterances that are often found together, such as question-answer, greeting-greeting, or offer-acceptance/refusal. The first part of the pair creates a expectation for the second part, thus forming a sequence.
- Repair mechanisms: These are strategies used by speakers to address problems in speaking, hearing, or understanding. Repair sequences involve the initiation of repair and its completion, forming a sequence within the conversation.
- Non-verbal cues: These include elements like gaze, gestures, and body posture, which can play a role in the sequential organization of conversation, such as signaling the end of a turn or the beginning of a new sequence.
All these elements contribute to the sequential structure of conversation, which is a key focus of conversation analysis.
By studying how these elements are ordered and related to each other, conversation analysts aim to uncover the underlying structure and organization of talk-in-interaction.
1. Turn-Taking
People in conversations generally respond to each other through a structured process known as turn-taking. It works like sharing: everyone gets a chance to speak.
Turn-taking keeps conversations orderly, with only one person talking at a time. It also lets people decide when to start and stop talking, so conversations flow smoothly with minimal overlap or awkward silence.
The foundational principle is simple. Conversation unfolds one speaker at a time. While that may seem obvious, CA provides a detailed framework for understanding how it is accomplished in practice.
1. Transition relevance places
Transition relevance places (TRPs) are points in conversation where speaker change becomes a possibility. These are often marked by the completion of a grammatical unit, a change in intonation, or a non-verbal cue.
At a TRP, the current speaker can either continue speaking or offer the floor to another participant.
There are two options here. The current speaker can directly address a specific participant, or use a more open-ended cue that allows anyone to self-select.
If the speaker doesn’t select the next speaker, other participants can self-select by starting to speak.
2. Turn allocation
Turn allocation techniques refer to the methods used by participants in a conversation to determine who speaks next.
These techniques are central to the organization of turn-taking, ensuring the smooth exchange of speaking turns with minimal overlapping or interruptions.
There are two primary types of turn allocation techniques: current speaker selects next and self-selection
2.1 Current speaker selects next
This technique involves the current speaker explicitly or implicitly choosing the next speaker.
Addressing a Question: One common method is directing a question to a specific participant, thereby selecting them to provide the answer.
For example, “Ben, do you want some?” explicitly selects Ben as the next speaker.
Affiliation with a first pair-part: A broader category of utterances, termed “first pair-parts,” also function as current-speaker-selects-next techniques. These include greetings, invitations and complaints.
Each initiates a specific type of adjacency pair, creating an expectation of a particular type of response from the selected recipient. Take this example: “Hey yuh took my chair by the way an’ I don’t think that was very nice.”
The first utterance is a complaint. It sets up an expectation of a denial or an account from the recipient in the next turn.
Lexical selection: The current speaker can also use specific words or phrases to indicate who should speak next, such as “never” or “ever” in a series of questions.
This alone does not guarantee they speak next.
For instance, if B answers A’s question and addresses their response to A, this doesn’t necessarily mean A is selected to speak next.
The context and subsequent actions of the participants will determine the actual turn allocation.
2.2 Self-selection
In contrast to the current speaker selecting the next speaker, self-selection occurs when a participant who was not selected initiates their turn at a transition relevance place (TRP).
- First Starter: The most prevalent form of self-selection is when a participant starts speaking first at a TRP. This is often marked by a slight gap or overlap with the previous speaker.
- Interruption: While less common, self-selection can also occur when a participant starts speaking before the current speaker has reached a TRP, effectively interrupting the ongoing turn.
This ordering of techniques ensures that conversations maintain their sequential organization and avoid multiple speakers talking simultaneously.
3. Context-dependent nature of turn-taking
A turn-taking system is considered “context-free” at its foundational level. This means the basic machinery for distributing turns operates consistently across diverse contexts, irrespective of who is speaking, what they are talking about, or their relationship.
This fundamental structure is what allows conversations to happen at all, ensuring that people mostly speak one at a time and transitions are generally smooth.
Context-Sensitivity: Turn allocation techniques aren’t arbitrary; they are influenced by the context of the conversation, the relationship between participants, and the social norms at play.
Understanding these techniques is crucial for researchers to analyze how conversations unfold, how participants manage their turns, and how meaning is co-constructed in interaction.
This context-free system becomes “context-sensitive” in its implementation. This means that while the underlying rules remain constant, how those rules are actually used to allocate turns can vary greatly depending on:
- Number of Participants: With more participants, the dynamics of turn distribution become more intricate, requiring adjustments to strategies for getting, keeping, or even relinquishing turns.
- Relationship Dynamics: The relationship between speakers (e.g., friends versus strangers, superiors versus subordinates) profoundly shapes turn-taking. Factors like relative status, familiarity, and existing power dynamics can influence who initiates, interrupts, or holds the floor for longer.
- Social Norms: Societal and cultural conventions heavily influence turn-taking. Norms dictate acceptable pauses, interruption etiquette, and how speakers signal turn transitions. These norms, often implicit, contribute to the “orderliness” of conversation within specific cultures or communities.
- Purpose of Conversation: The goal of a conversation (e.g., casual chat, formal debate, institutional interaction) influences turn-taking. For example, debates have pre-allocated turns, unlike informal conversations where turn allocation is more fluid.
- Turn Content as a Factor: Importantly, while the turn-taking system is indifferent to the content of turns, this doesn’t mean content is irrelevant. What speakers say in their turns provides the context for subsequent turns. For example, a question (content) projects an answer as the relevant next action, shaping turn allocation.
The underlying system provides the framework, in essence. But the actual dance of conversation, with its nuances and adjustments, emerges from the interplay of this context-free base with the ever-changing dynamics of context, relationships and social expectation.
The Turn-Taking Systematics Study
Everything above rests on one founding paper: Sacks, Schegloff and Jefferson’s (1974) “A simplest systematics for the organization of turn-taking for conversation.”
It was the first study to show mundane talk is a rule-governed social institution, not an unstructured habit.
Aim: To give a formal, empirically grounded account of how speaker change is achieved, rather than assuming ordinary talk is chaotic or run to a fixed order.
Method: The team worked inductively from real recordings. This included years of naturally occurring talk, such as phone calls and group-therapy sessions, examined for recurring patterns via “unmotivated looking” and checked against deviant cases rather than discarding them.
Results: Turn-taking rests on turn-constructional units. Their completion projects a transition relevance place, and an ordered rule set then governs who speaks next; overlaps between turns are systematically minimised.
Conclusion: Turn-taking is locally managed, party-administered, and both context-free and context-sensitive. The same machinery has since been confirmed in the classroom (McHoul, 1978) and the courtroom (Atkinson & Drew, 1979).
2. Adjacency Pairs
Adjacency pairs are pairs of turns in conversation that are closely related to each other, and they are the fundamental unit of sequencing.
A pair consists of two utterances by different speakers, in which the first part makes the second part conditionally relevant. The term matters.
Schegloff (1968) introduced this concept in his analysis of how telephone calls are opened.
That means the first part sets an expectation for a specific type of response in the second. Examples include question-answer, invitation-acceptance/decline, and greeting-greeting.
For example, a greeting is typically followed by another greeting, a farewell by a farewell, and a question by an answer. These pairs are considered the basic building blocks of conversational sequences.
Here are some key features of adjacency pairs:
- They consist of two turns made by different speakers.
- These turns are typically adjacent to each other, meaning there is no intervening talk between them.
- The turns are ordered so that the first pair part (FPP) always comes before the second pair part (SPP). For example, a question always precedes its answer.
- The turns are differentiated into pair types where the FPP makes a particular type of SPP relevant. Examples of adjacency pair types include:
- question – answer
- greeting – greeting
- summons – answer
- telling – accept
The concept of conditional relevance is central to understanding adjacency pairs. This principle states that the first part of an adjacency pair establishes an expectation for a particular type of second part.
If the expected second part does not occur, it is considered “noticeably absent.” This absence becomes significant for analyzing the conversation.
Adjacency pairs can be expanded in several ways:
- Preface: A preface expands an adjacency pair with pre-expansions (like disclaimers), which are sequences that prepare for upcoming talk and project that further talk will follow.
- Extension: An extension expands an adjacency pair with post-expansions (like assessments), which are sequences that follow and extend the base adjacency pair.
- Insertion: Insert expansions, which are nested between the first and second pair parts of an adjacency pair, may be used to clarify information before responding to the first pair part
Adjacency pairs are a fundamental aspect of conversation analysis, but they are not the only way conversations are organized.
3. Repair Mechanisms
Repair in conversation analysis refers to the processes by which speakers address problems that arise in talk. These problems can be related to hearing, production, or understanding.
Rather than being a negative phenomenon, repair is a natural and essential self-regulating device crucial for maintaining coherence in conversation.
Types of repair
Repair mechanisms can be categorized based on who initiates the repair (the speaker or the recipient of the problematic utterance) and who carries out the repair. This results in four primary types of repair:
- Self-initiated self-repair: The speaker of the problematic utterance both identifies and resolves the issue.
- Self-initiated other-repair: The speaker identifies a problem, but the recipient provides the solution.
- Other-initiated self-repair: The recipient of the problematic utterance points out the issue, and the original speaker resolves it.
- Other-initiated other-repair: The recipient both identifies and resolves the problem in the speaker’s utterance.
Preference for self-repair
While both speakers and recipients can initiate repair, there’s a general preference for self-initiated repair over other-initiated repair.
This suggests that the conversational system prioritizes speakers taking responsibility for their own utterances.
Other-initiated repair often takes the form of signaling a problem, prompting the original speaker to carry out the repair (other-initiated self-repair).
This highlights that even when others initiate repair, the system is often structured to ultimately facilitate self-repair.
Repair positions and turn-taking
Repair is also closely tied to turn-taking systems in conversation. The timing of repair initiation, relative to the problematic utterance, shapes how repair unfolds:
- Same-turn repair: The speaker initiates repair within the same turn as the problematic utterance, often using non-lexical cues like cut-offs, sound stretches, or pauses.
- Transition space repair: Repair happens in the brief silence between turns, after the problematic turn.
- Second position repair: Repair occurs in the turn immediately following the one containing the problem. This often involves the recipient initiating the repair.
- Third position repair: Repair takes place in the third turn after the problematic utterance.
- Fourth position repair: While less common, repair can occur in the fourth turn, usually involving the recipient addressing a persistent issue.
The sequential organization of repair, with opportunities for self-repair preceding those for other-repair, further emphasizes the system’s design favoring self-correction.
Repair in online communication
Repair in online communication, while sharing similarities with face-to-face interaction, also exhibits differences due to the nature of the medium.
One notable difference is the absence of same-turn repair in online interactions, as any corrections made during typing wouldn’t be visible to others.
Additionally, online communication might see a weakening of the preference for self-repair, as the recipient might resolve the issue more efficiently in some cases.
4. Non-Verbal Cues In Communication
Non-verbal cues are central to human communication, and though communication is possible without them (as in telephone calls), that does not make them peripheral to the process.
Human communication is inherently multimodal, meaning it uses all available modes to convey information between speakers and recipients. These can include less easily recorded modes like smell or taste.
How non-verbal cues shape meaning
Non-verbal elements like laughter, smiling, intonation, and stress act as contextualization cues. They work alongside language to shape how words are understood in a given interaction.
This means that while they don’t inherently encode meaning on their own, they influence how spoken words are interpreted.
For example, a statement can be interpreted literally if accompanied by laughter or a smile, while intonation and stress can convey sarcasm or a negative response.
Gaze in conversation
Gaze, as an act of seeing and a communicative act, plays a significant role in social interaction. It signals what a participant is attending to and can be used to solicit a response, even without explicit verbal prompting.
If a speaker has not received a response to their talk, for instance, they might use gaze to elicit one. That response might be verbal or a gesture.
Gesture as Communication
Gestures, which convey meaning through bodily action, are not incidental but a core part of interaction. They have a central communicative function that contributes to the overall meaning-making in conversation. Some of the roles gestures play within interactions include:
- Turn Allocation: Gestures can be used to gain the floor in a conversation, acting as a way to signal a desire to speak.
- Turn-Taking Organization: They contribute to the non-verbal aspects of turn-taking, such as providing cues for turn completion.
- Replacing Linguistic Forms: Gestures can stand alone as complete turns, replacing verbal communication entirely. For instance, nodding or shaking one’s head can convey agreement or disagreement, a wave acts as a greeting or farewell, and redirecting gaze can be a response to being addressed.
Integrating gesture, gaze, and talk
In conversation, gesture, gaze and talk work in a coordinated way to construct meaning. Gestures can introduce non-present entities into a conversation, much as pointing incorporates physically present objects.
A speaker might gesture toward a specific location on a screen while saying the word “problem.” Both are deictic.
They work together to direct attention to a spatial location and establish shared focus. That matters.
The gesture is not just a repeat of the spoken word, though. The word relies on the gesture for its full meaning: without it, the location might stay unclear.
Goodwin (2000) calls this integration a laminated action: talk, gesture and gaze are layered together, and no single layer alone carries the meaning participants build.
Non-verbal cues in conversation analysis
Conversation analysis (CA) emphasizes the significance of non-verbal cues in understanding social interaction.
Analyzing elements like speaking speed and intonation provides valuable context for comprehending the nuances of social interaction.
For instance, a speaker’s confidence level when answering a question can be inferred from their intonation and pauses.
Hesitations or pauses might suggest uncertainty as the speaker searches for the right words. Conversely, emphasizing certain words can convey authority and expertise.
By considering these non-verbal cues, CA provides a richer understanding of the meaning conveyed in interactions beyond the literal words spoken.
Steps for Conducting CA
- Data Collection: Collect data using audio or video recordings of naturally occurring interactions.
- Transcription: Transcribe the recordings in detail, using a system like the Jeffersonian transcription system to capture pauses, intonation, and other non-verbal cues.
- Unmotivated Looking: Listen to the recordings multiple times without any pre-existing theories in mind.
- Identify Phenomena: Identify recurring patterns in the data, such as turn-taking, repair strategies, and the use of specific words or phrases.
- Analyze the Data: Analyze the identified phenomena. For instance, how do speakers use pauses to manage turn-taking or how do they repair misunderstandings?
- Develop an Analysis: Develop a clear and concise analysis, focusing on the sequential organization of the talk. For example, how does a speaker’s turn relate to the previous turn?
- Contextualize the Analysis: Consider the context of the interaction. What are the social and cultural norms that might be influencing the interaction? What are the individual differences between the speakers?
Step 1: Data Collection
Data Collection in CA focuses on gathering recordings of these naturally occurring conversations. This could be conversations between friends, family members, or even strangers. The idea is to capture how people actually talk, not how we think they talk.
The goal is to gather data that accurately reflects real-world conversations for analysis.
Example
A researcher studying how people apologize, for example, might collect recordings of conversations where apologies occur naturally, such as between friends who have had a disagreement. They would not ask friends to stage an argument and apologize. That would not reflect how people genuinely interact.
Instead, they might identify situations where apologies are likely: after a friend forgets a promise, or accidentally says something hurtful. They could then ask whether those friends would be willing to be recorded in such situations.
The goal is authenticity. Capturing genuine apologies in their natural context gives insight into how people use language to repair relationships and navigate social dynamics.
Recordings should capture as much detail as possible, including pauses, intonation and word stress, which requires audio or video recording equipment.
The Observer’s Paradox
CA researchers acknowledge that recording interactions might influence how naturally participants behave. This is the observer’s paradox.
The term was coined by the sociolinguist William Labov (1972). It names a general problem: the data researchers most want, how people talk when not being observed, can typically only be obtained by observing them.
That risks changing the very behaviour being studied. Ideally, CA seeks to understand how people interact when they are not being observed.
Researchers try to minimize the impact of recording with unobtrusive methods, such as an “absent observer” where only a recording device is present, not the researcher.
Even so, the possibility remains that participants’ behavior is influenced by the research process.
Participants’ Awareness of Recording
There is some evidence that recording devices do not always significantly affect interaction. This isn’t settled. Speer and Hutchby (2003) argue that participants’ reactions to being recorded can themselves be analyzed within the context of the interaction.
The impact varies by context.
Importantly, ethical considerations in CA research require that participants are always aware they are being recorded and consent to it. While this awareness might impact the naturalness of their interaction, ethical research practices prioritize informed consent.
Step 2: Transcription
Transcribing talk in conversation analysis involves more than simply recording the words spoken. Conversation analysts need to know “how it was said” in addition to “what has been said”. The transcript should capture features like pauses, intonation, stress, and overlapping speech.
Conversation analysis employs a meticulous transcription system developed by Gail Jefferson. This system is designed to capture the nuances of naturally occurring talk for in-depth analysis.
Here are key aspects of this transcription technique:
- Detailed Representation: This transcription method goes beyond just words, aiming to represent pauses, overlaps, intonation, and even non-verbal aspects like laughter and breathing. This intricate approach allows researchers to see the transient, complex nature of talk in a static, analyzable format.
- Specialized Symbols: The Jeffersonian transcription system utilizes a unique set of symbols to denote these conversational elements, like micropauses, overlapping speech, intonation, and more. This system helps researchers capture the precise delivery of talk, including pace, overlapping talk, and intonation.
- Iterative Refinement: Creating a conversation analysis transcript is an iterative process. Researchers often start with basic elements like words and pauses, then layer in more complex information like intonation, stress, and overlapping speech. This layered approach to transcription ensures the capture of subtle details for a comprehensive analysis.
- Importance of Context: Transcripts alone are not sufficient for analysis; they are always used alongside recordings. Researchers constantly revisit and refine their transcripts based on repeated listening to the recordings. Capturing non-verbal cues like gaze, gestures, and object interaction requires video recordings, adding further complexity to the transcription process.
Recording in natural settings with audio and video has real advantages over intuition or invented sentences, despite the limits of capturing every detail. Transcription itself is selective.
It is shaped by the researcher’s own analytical goals, which helps identify recurring patterns and subtle nuances in social interaction.
Automatic transcription software is of limited use here. It is adequate for basic transcription but struggles with overlapping talk, and it cannot capture the fine detail the analysis depends on.
Specialized Symbols
The Jefferson Transcription System meticulously details the nuances of spoken interaction, going beyond mere words to include pauses, overlaps, intonation, and even non-verbal elements like laughter.
This system is not simply about accurate documentation; it provides a structured framework for analyzing the complexities of naturally occurring talk.
Turn-taking
Turn-taking is the systematic allocation of opportunities to talk, and the regulation of the size of those opportunities.
To account for turn-taking dynamics, transcripts aim to capture the details of how turns are taken in talk-in-interaction.
These details include the precise points at which turns begin and end, including overlaps, gaps, pauses, and audible breathing.
- Overlapping Speech: Square brackets ([]) precisely mark the points where simultaneous speech occurs, capturing the intricacies of interruptions, simultaneous starts, and turn-taking competition.
- Contiguous Utterances: An equal sign (=) signifies a swift transition between consecutive utterances without a discernible pause, highlighting the rapid flow of speech.
- Breathing: The symbol “.hhh” indicates an in-breath, further adding to the transcription’s representation of the speaker’s delivery.
- Pauses: The duration of silences is precisely measured, typically in tenths of a second, using numerals within parentheses. This precision highlights the interactional significance even brief pauses can hold, as demonstrated by research showing the impact of pauses as short as two or three-tenths of a second.
Speech Delivery
To account for the characteristics of speech delivery, transcripts mark noticeable features of stress, enunciation, intonation, and pitch.
For example, if a speaker noticeably extends a word, colons are inserted into the word at the point of extension. The longer the audible extension, the more colons are inserted.
- Intonation: Punctuation marks are repurposed to denote intonation: a period (.) for a falling tone, a question mark (?) for a rising tone, and a comma (,) for a non-final, flat tone.
- Stress: Underlining beneath a word signifies emphasis or stress, drawing attention to words given prominence in spoken delivery.
- Sound Stretching: Colons (:) visually represent the lengthening of a sound. The number of colons corresponds to the duration of the extension, providing a visual representation of drawn-out pronunciation.
- Inaudible Speech: Parentheses with empty space ( ( ) ) are used to denote instances where speech is indistinguishable, acknowledging the limits of transcription while maintaining the sequential flow of the conversation.
Beyond Words: Capturing Non-Verbal Communication
While the Jefferson Transcription System excels in capturing the nuances of spoken language, it also acknowledges the importance of non-verbal elements in interaction.
Researchers have expanded the system to encompass visual information, especially in video-recorded data.
For example, symbols like paired asterisks (*) or carets (^) can denote gestures made by different speakers.
Context matters here too. Descriptive annotations within double parentheses ( (()) ) provide context about actions, such as a car turning a corner. This enriches understanding of the interaction’s setting and its potential influence on the dialogue.
Step 3: Unmotivated Looking
Listen to the recordings multiple times without any pre-existing theories in mind.
Listening to the recordings
During the “Unmotivated Looking” stage of conversation analysis (CA), the researcher repeatedly listens to the same recordings.
This process aims to understand what transpires in the data without imposing preconceived theories or expectations.
The focus is on uncovering naturally occurring patterns and structures within the conversation.
Openness to discovery
Unmotivated looking encourages the analyst to be receptive to discovering unexpected phenomena in the data.
Rather than searching for specific pre-identified elements, the researcher maintains an open mind, allowing the data to guide their observations.
This approach helps in identifying subtle but significant aspects of social interaction that might otherwise be overlooked.
Noticing and identifying actions
The process involves carefully attending to the details of the talk, including seemingly insignificant features. The goal is to understand the actions being performed through language.
For example, a researcher might notice a particular phrase and try to identify its effect on the subsequent conversation.
This can involve identifying how participants use specific practices to achieve communicative goals, like making a request or offering an assessment.
Challenges and considerations
Complete neutrality is difficult, though: prior knowledge and research interests inevitably shape perception. The process means balancing openness to new discoveries against the existing body of knowledge in CA.
Despite these challenges, unmotivated looking remains a fundamental principle of the method.
Step 4: Identify Phenomena
The focus shifts to identifying recurring patterns in the data, such as turn-taking, repair strategies, and the use of specific words or phrases.
These patterns can manifest in various ways, including:
- Turn-Taking: This involves analyzing how speakers alternate turns in a conversation, examining elements like turn allocation and speaker selection. For instance, identifying instances where the current speaker selects the next speaker or examining how overlapping talk is managed.
- Repair Strategies: This entails studying how participants address and resolve communication breakdowns or misunderstandings. Examples include noting where repair work occurs, identifying the type of repair (e.g., self-initiated self-repair), and analyzing how participants construct and respond to repair attempts.
- Use of Specific Words or Phrases: This involves recognizing recurring linguistic features, such as particular words, phrases, or grammatical structures that hold significance in the data. This can include examining the use of explicit repair devices (e.g., “Excuse me”) or identifying specific formats used to perform particular actions (e.g., “You should X” for requests).
The goal is to move beyond individual instances and identify patterns that reveal how participants understand and navigate social interaction.
For example, analyzing instances of third-position repair, a pattern where a speaker clarifies their prior utterance after the recipient’s turn shows a problem in understanding.
Recognizing these recurring patterns helps researchers develop a deeper understanding of the practices and conventions governing conversation.
Step 5: Analyze the Data
The next step in analyzing conversational data is to examine each case in the collection by analyzing : activity, participation, position, composition, and action of the conversation.
Analyzing each of these aspects creates an understanding of how the interaction functions line by line.
- Activity: This refers to what participants are doing together through their interaction. When examining activity, some questions to consider are:
- What are the circumstances of the interaction?
- Do the participants share a common goal, environment, or communication medium?
- Is there a goal, or is it more loosely organized?
- Are certain actions done at certain times, in a certain order, or by certain participants?
- Participation: This refers to the roles participants occupy throughout the interaction. When examining participation, some questions to consider are:
- What roles do the participants occupy generally (e.g., speaker vs. the person who just finished speaking).
- What roles do participants occupy turn-by-turn (e.g., speaker and recipient)?
- What roles do participants occupy within a sequence (e.g., someone initiating a repair)?
- What roles do participants occupy within the specific occasion (e.g., caller and receiver)?
- How do the participants navigate and change roles?
- Position and composition: Where the conduct sits in the sequence, and how it is built. A characterization of action should come only after an adequate analysis of sequence structure and turn construction. This is the core distinction.
- Action: This refers to what the talk and conduct accomplish in the interaction. The location of the conduct within the conversation and how it is formatted make up the action.
The goal of analyzing each case in the collection line by line is to be able to produce a comprehensive analysis of each conversation.
As you examine the conversations, take notes and modify the formal description of the phenomenon as needed.
Step 6: Develop an Analysis
The final step in analyzing conversational data is to develop a formal account of the phenomenon.
The criteria used to identify the phenomenon and its boundaries are key to this account, along with an analysis of variations across the entire collection of conversations.
The account should describe the phenomenon’s structure, including the linguistic forms and social actions involved. It should explain how it functions, the conditions in which variations arise, and the interactional problem it addresses.
Example of developing an analysis
Take a particular type of question-response sequence. The analysis might examine how specific linguistic forms, such as hedges, relate to participants’ understanding of each other’s knowledge of the topic.
The word “sounds” in the assessment “That sounds interesting” may indicate that the speaker assumes the recipient has limited knowledge of the object being assessed.
A formal account might propose that speakers use hedges like “sounds” to convey a lack of certainty about an assessment. The logic is simple.
This fits when they believe the recipient has no direct experience of the object.
That account needs grounding in evidence from the collection, showing that speakers consistently use hedges in these contexts.
It might then explain that hedging in assessments serves to manage social epistemics. That is the mechanism.
This means who has primary rights to know something, and who has only secondary access to it (Heritage & Raymond, 2005).
Hedges acknowledge the recipient’s limited knowledge and help avoid potential challenges or disagreement, connecting this small linguistic choice to a broader social function.
Step 7: Contextualize the Analysis
Consider the context of the interaction.
When analyzing conversations, it is important to consider the social and cultural norms that might be influencing the interaction.
For example, the way people take turns speaking or the types of speech acts that are considered appropriate can vary depending on the culture.
Additionally, the social relationships between the speakers, such as whether they are friends, family members, or strangers, can also influence how they interact.
Conversation analysis (CA) focuses on analyzing how participants in an interaction understand and shape the interaction. This is key. It does not impose external assumptions about the influence of social categories or relationships.
Next-turn proof procedure
The next-turn proof procedure was formulated by Heritage (1984) from Sacks, Schegloff and Jefferson’s work. It is a basic tool in CA.
A turn is analyzed as evidence of its speaker’s understanding of the prior turn. It rests on a simple fact.
The turn-taking system requires speakers to display their understanding of the prior turn, in order to produce a relevant next turn. Consider a question.
When a speaker asks one, the next speaker is expected to answer. By answering, they display their understanding that the prior turn was a question, and it is that display the analyst examines.
The procedure is valuable for a simple reason. It lets analysts see how participants make sense of each other’s turns.
This keeps the analysis grounded in participants’ own understanding, rather than the analyst’s assumptions.
Here are some questions to consider when contextualizing conversational analysis:
- What are the cultural backgrounds of the speakers?
- What is the relationship between the speakers? (e.g., friends, family, colleagues, strangers)
- What is the setting of the interaction? (e.g., formal or informal, public or private)
- What is the purpose of the interaction? (e.g., to exchange information, to build relationships, to accomplish a task)
By considering these contextual factors, researchers can gain a more complete understanding of the interaction and how the participants are using language to achieve their goals.
Applications of Conversation Analysis
CA’s insistence on naturally occurring data has taken it well beyond linguistics and sociology. It now reaches any setting where an institution’s outcomes are actually accomplished turn by turn.
Clinical and Therapeutic Talk
CA’s naturalistic commitment extends even into the therapy room. Applying it to psychotherapy means treating the therapy session itself as the object of analysis, not just what the client reports about their life outside it.
How a therapist’s next turn is designed matters. So does where an interpretation is placed in the sequence.
How clients respond to it matters too: accepting it, resisting it, or working to reframe it. These become observable, describable events.
They are no longer just inferred internal processes (Peräkylä, Antaki, Vehviläinen, & Leudar, 2008). This reframes therapeutic technique as something researchers can study empirically.
It uses the same turn-taking, repair and adjacency-pair machinery documented in ordinary talk. That differs from relying only on client self-report or before-and-after outcome measures taken at the start and end of treatment.
Medical Communication
Doctor-patient consultations are among the most heavily studied institutional settings in CA. That is because the outcome is visibly built turn by turn, rather than simply reported afterwards.
This includes a diagnosis accepted, or a prescription written. Maynard and Heritage (2005) show that physicians typically retain control of the topic and direction of a consultation, even where it sounds conversational.
This echoes the asymmetries CA documents elsewhere. Stivers (2007) demonstrates the practical stakes in concrete terms.
Consider a physician’s diagnosis. How they frame it just before a treatment recommendation, as a “problem” versus a “no problem,” measurably shapes whether parents accept a decision not to prescribe antibiotics.
Together, these studies show CA moving from describing institutional talk to identifying the specific turn-design choices that predict a real clinical outcome.
Education
Classroom talk has its own, more constrained turn-taking economy than ordinary conversation. McHoul (1978) shows that turn allocation in formal classroom talk is overwhelmingly controlled by the teacher.
This differs from the self-selection and current-speaker-selects-next rules that organise mundane conversation. Students typically speak only when nominated.
The teacher alone retains the right to allocate, interrupt and close down turns, in a way no single party has in ordinary talk. The asymmetry is built into the structure of the lesson itself, not just the teacher’s personal style.
This gives researchers and teacher-trainers a precise, turn-by-turn vocabulary for describing classroom participation. That is more useful than a vague label.
A lesson called “discussion-based” can still turn out, turn by turn, to be teacher-controlled throughout. CA is what lets a researcher tell the difference.
Forensic and Legal Settings
Courtroom talk is a limiting case of institutional asymmetry. Question-answer sequences are tightly pre-allocated, and witnesses are ordinarily restricted to answering rather than initiating.
Atkinson and Drew’s (1979) foundational study of courtroom interaction shows how the precise design of a question constrains what a witness can accountably say next.
Question design is not neutral. It includes a question’s presuppositions and the response options it makes available.
The outcome of an exchange is shaped as much by question design as by the “facts” a witness reports. The stakes are real.
Control over turn-taking in an adversarial setting is itself a form of control over what can be said. This power extends well beyond the courtroom, into any institutional setting with pre-allocated turns.
Human-Computer Interaction and Conversational AI
CA’s turn-taking and repair apparatus was developed from human-to-human talk. It has more recently been applied to how people talk to voice assistants and other conversational technologies.
Porcheron, Fischer, Reeves and Sharples (2018) recorded Amazon Echo devices in participants’ own homes for a month. Their study drew explicitly on ethnomethodology and conversation analysis.
It documents how households practically fit voice-interface use into their existing, ongoing social interactions. That is different from treating it as an isolated human-machine exchange. The point generalises.
The same turn-taking machinery that organises a dinner-table conversation, in other words, also shapes how a family talks to a smart speaker. The parallel is real, and it is discussed further, alongside a second recent study, in the Critical Evaluation section below.
Tips for Conducting Conversation Analysis
- Focus on naturally occurring interactions: Use actual talk in context, such as recordings from everyday conversations or institutional settings. Conversation Analysis (CA) emphasizes studying real-world language to understand how communication works in its natural environment. Avoid using manipulated or artificial data.
- Prioritize detailed transcription: Capture the nuances of spoken interaction, including pauses, intonation, word stress, and non-verbal cues. Transcribe repetitions, grammatical errors, and other speech features that might be relevant for understanding the interaction. The Jeffersonian transcription system is commonly used in CA research to denote these details.
- Use video recording for comprehensive data: Video recordings capture non-verbal cues like gestures, body language, and shared visual context that might be missed in audio-only recordings. While audio recording has its advantages (e.g., less intrusive, easier to transcribe), video provides a more complete record of the interaction for analysis.
- Engage in ‘unmotivated looking’: Listen to the recordings repeatedly without preconceived notions to identify interesting or puzzling patterns. This inductive approach helps uncover recurring patterns and discover new insights directly from the data.
- Analyze sequential organization: Examine how turns are taken, actions are coordinated, and meaning is co-constructed within the conversation flow. The order and placement of utterances are crucial for understanding their meaning and function.
- Consider the context: Acknowledge the influence of social and cultural factors, individual differences, and the specific situation on the interaction. CA recognizes that communication is shaped by its context and avoids assuming universal rules of interaction.
- Avoid imposing pre-theorized frameworks: Let the data guide the analysis and develop theoretical insights inductively. While CA research might incorporate existing theories, the primary focus is to understand interactional patterns based on evidence from the data itself.
- Use clear and consistent notation: When presenting findings, employ a widely recognized transcription system and explain any deviations or specific notations. Consistent notation ensures that others can understand and interpret the analysis.
- Thoroughly explain the data and analysis: Present findings in a way that allows the audience to understand the context, the reasoning behind the analysis, and the significance of the observed patterns. Provide sufficient detail to support the claims and enable others to follow the analytical process.
Critical Evaluation of Conversation Analysis
Like any method, CA has real strengths and real limits. These have generated genuine, still-unresolved methodological debate.
Strengths of the Naturalistic Approach
CA’s founding commitment is to analyse recordings of talk that would have happened whether or not a researcher was present. This is what gives its findings their claim to ecological validity.
Sacks, Schegloff and Jefferson (1974) built the turn-taking systematics inductively from such recordings, rather than from invented or role-played examples. The goal is simple.
CA aims to describe the methods people actually use, which staged or hypothetical talk cannot straightforwardly reveal. This is widely regarded as CA’s central strength relative to methods that rely on interviews or self-report.
Self-report captures what participants say they do, not what a recording shows they actually do. That distinction matters.
CA also avoids imposing the analyst’s own categories on the data. An analyst cannot simply assert that talk is “about power” or “resistance.”
They must demonstrate it turn by turn. This is the next-turn proof procedure (Heritage, 1984): showing that participants themselves are oriented to it that way.
Key Limitations
Labour-intensive transcription and small collections limit CA’s scale and generalisability. Jeffersonian transcription (Jefferson, 2004) is deliberately exhaustive: pauses are timed to tenths of a second, and overlap is marked to the syllable.
Producing even a short stretch of it to a usable standard takes far longer than the interaction itself lasts. That caps how much data any single study can cover.
Sacks, Schegloff and Jefferson’s (1974) original systematics was itself built from a comparatively narrow, largely American-English corpus. CA’s claim to describe “context-free” machinery therefore rests on later studies, not the founding corpus alone.
McHoul’s (1978) classroom data and Atkinson and Drew’s (1979) courtroom data are two such replications, in different institutional settings. That is a real constraint.
A single collection of recordings, however carefully analysed, cannot by itself establish that a pattern generalises beyond it.
The Discourse-Analysis Critique
Not every critic accepts that CA’s restraint about context is a virtue. Schegloff (1997) defends CA’s own position.
He argues that discourse-analytic approaches which read talk as an expression of power risk imposing those categories from outside. That is true unless participants can be shown, in the sequential detail of the talk itself, to be actually orienting to them.
Otherwise, the “power” being described is the analyst’s interpretation, not a demonstrated feature of the interaction. Billig (1999) rebuts this directly.
He argues that CA’s own technical, apparently neutral categories are themselves interpretive choices. They can quietly naturalise existing power relations, rather than escaping the problem Schegloff raises.
Refusing to name power in advance, on this view, is not the same as having no position on it. This remains a live, unresolved disagreement between the two traditions. Neither side has settled it.
Contemporary Research
A consistent theme in post-2015 conversation-analytic research is that the classic turn-taking machinery does not transfer unchanged to newer technology. It was developed from co-present and telephone talk.
The strongest evidence for this comes from systematic synthesis. Seuren, Ilomäki, Dalmaijer, Shaw and Stommel (2024) conducted a systematic review of conversation-analytic telehealth research.
Aim: To synthesise the previously unreviewed body of CA research on remote healthcare interaction.
Method: The authors searched relevant databases to January 2022, updated in April 2023. The net was wide.
They screened for studies using CA on naturally occurring telehealth interactions between a clinician and a patient.
Two reviewers independently screened the first 200 records, with 96% agreement; 41 articles met the criteria. The results were clear.
Findings: Whether participants can see each other measurably affects turn-taking. Video-call latency causes people to fall silent or talk over each other.
A silence between text messages, meanwhile, means something different from a silence in a video call. Text is not just delayed speech.
Clinicians must also do visible interactional work to sustain patient engagement in text-based care.
Conclusion: “Remote healthcare encounters are not defective forms of in-person healthcare, but different forms of care.” Each has its own interactional organisation that participants competently manage.
It is not simply a degraded version of co-present consultations. Porcheron, Fischer, Reeves and Sharples’s (2018) month-long study of Amazon Echo devices in participants’ homes points the same way.
It found that voice-assistant use is threaded into a household’s ongoing social interaction, rather than kept separate from it. The same underlying point, in other words, shows up even in a non-clinical, human-AI context.
Key Takeaways
- Definition: CA studies how talk is organised as an orderly, rule-governed activity, not just what is said.
- Origins: Developed by Harvey Sacks, Emanuel Schegloff and Gail Jefferson at UCLA in the 1960s-70s, rooted in Garfinkel’s ethnomethodology.
- Turn-Taking: Conversation is locally managed one speaker at a time, via transition relevance places and an ordered set of allocation rules.
- Repair: Four repair types exist, with a strong preference for self-repair over other-repair.
- Applications: Used in clinical and therapeutic talk, medical consultations, classrooms, courtrooms, and now voice-assistant interaction.
- Limitations: Jeffersonian transcription is labour-intensive, and the founding corpus was narrow (American English, telephone calls, group therapy).
- Evidence caveat: Seuren et al. (2024) show CA’s turn-taking machinery needs rethinking, not just reapplying, for telehealth and text-based care.
Further Information
Atkinson, J. M., & Drew, P. (1979). Order in court: The organisation of verbal interaction in judicial settings. Macmillan.
Billig, M. (1999). Whose terms? Whose ordinariness? Rhetoric and ideology in conversation analysis. Discourse & Society, 10(4), 543–558. https://doi.org/10.1177/0957926599010004005
Goodwin, C. (2000). Action and embodiment within situated human interaction. Journal of Pragmatics, 32(10), 1489–1522. https://doi.org/10.1016/S0378-2166(99)00096-X
Heritage, J. (1984). Garfinkel and ethnomethodology. Polity Press.
Jefferson, G. (2004). Glossary of transcript symbols with an introduction. In G. H. Lerner (Ed.), Conversation analysis: Studies from the first generation (pp. 13–31). John Benjamins. https://doi.org/10.1075/pbns.125.02jef
Maynard, D. W., & Heritage, J. (2005). Conversation analysis, doctor–patient interaction and medical communication. Medical Education, 39(4), 428–435.
Peräkylä, A., Antaki, C., Vehviläinen, S., & Leudar, I. (Eds.). (2008). Conversation analysis and psychotherapy. Cambridge University Press.
Sacks, H., Schegloff, E. A., & Jefferson, G. (1978). A simplest systematics for the organization of turn taking for conversation. In Studies in the organization of conversational interaction (pp. 7-55). Academic Press.
Schegloff, E. A. (1987). Analyzing single episodes of interaction: An exercise in conversation analysis. Social psychology quarterly, 101-114.
Schegloff, E. A. (1992). Repair after next turn: The last structurally provided defense of intersubjectivity in conversation. American journal of sociology, 97(5), 1295-1345.
Schegloff, E. A. (1993). Reflections on quantification in the study of conversation. Research on Language and Social Interaction, 26(1), 99–128.
Schegloff, E. A. (2002). 18 Beginnings in the telephone. Perpetual contact: Mobile communication, private talk, public performance, 284.
Schegloff, E. A. (2007). Categories in action: Personreference and membership categorization. Discourse
Studies, 9(4), 433–461.
Schegloff, E. A. (2007). Sequence organization in interaction: A primer in conversation analysis I (Vol. 1). Cambridge university press.
Schegloff, E. A. (2007). A tutorial on membership categorization. Journal of pragmatics, 39(3), 462-482.
Schegloff, E. A., & Sacks, H. (1973). Opening up closings.
Speer, S. A., & Hutchby, I. (2003). From ethics to analytics: Aspects of participants’ orientations to the presence and relevance of recording devices. Sociology, 37(2), 315-337.
Stivers, T. (2007). Prescribing under pressure: Parent-physician conversations and antibiotics. Oxford University Press.