The observation method in psychology involves directly and systematically recording measurable behaviors, actions, and responses in natural or contrived settings. The researcher does not manipulate what is observed.
Used to describe phenomena, generate hypotheses, or validate self-reports, psychological observation can be either controlled or naturalistic, with varying degrees of structure imposed by the researcher.
There are different types of observational methods, and distinctions need to be made between:
1. Controlled Observations
2. Naturalistic Observations
3. Participant Observations
Observations can also be overt or covert. In overt/disclosed studies, participants know they are being watched. In covert/undisclosed studies, the researcher’s real identity is kept secret from the research subjects, and the researcher acts as a genuine member of the group.
In general, conducting observational research is relatively inexpensive, but it remains highly time-consuming and resource-intensive in data processing and analysis.
The considerable investments needed in terms of coder time commitments for training, maintaining reliability, preventing drift, and coding complex dynamic interactions place practical barriers on observers with limited resources.
Controlled Observation
Controlled observation is a research method for studying behavior in a carefully controlled and structured environment.
The researcher sets specific conditions, variables, and procedures to systematically observe and measure behavior, allowing for greater control and comparison of different conditions or groups.
The researcher decides where the observation will occur, at what time, with which participants, and in what circumstances, and uses a standardized procedure. Participants are randomly allocated to each independent variable group.
Detailed description of every behavior observed is often impractical. Researchers instead code behavior against a previously agreed scale, using a behavior schedule to conduct a structured observation.
The researcher classifies each behavior into distinct categories, often using numbers or letters. A scale can also capture how intense the behavior was, not just whether it occurred.
These codes are quick to count and convert into statistics.
For example, Mary Ainsworth used a behavior schedule to study how infants responded to brief periods of separation from their mothers. During the Strange Situation procedure, the infant’s interaction behaviors directed toward the mother were measured, e.g.,
- Proximity and contact-seeking
- Contact maintaining
- Avoidance of proximity and contact
- Resistance to contact and comforting
The observer scored behavior in 15-second intervals. Each interval was rated for intensity on a 1–7 scale.

Some studies observe participants through a two-way mirror, or film them secretly. Albert Bandura used this method to study aggression in children (the Bobo doll studies).
A lot of research has been carried out in sleep laboratories as well. Here, electrodes are attached to the scalp of participants. What is observed are the changes in electrical activity in the brain during sleep (the machine is called an EEG).
Controlled observations are usually overt as the researcher explains the research aim to the group so the participants know they are being observed.
Controlled observations are also usually non-participant as the researcher avoids direct contact with the group and keeps a distance (e.g., observing behind a two-way mirror).
Strengths
- Replicable: The same observation schedule can be reused by other researchers, making it easy to test for reliability.
- Quick to Analyze: The data is quantitative, so it is faster to analyze than naturalistic observation data.
- Large Samples: Observations are quick to conduct, so a large sample can be obtained, making findings more representative and generalizable.
Limitations
- Lacks Validity: The Hawthorne effect and demand characteristics mean participants may act differently once they know they are being watched, threatening validity.
Naturalistic Observation
Naturalistic observation is a research method in which the researcher studies behavior in its natural setting without intervention or manipulation.
It involves observing and recording behavior as it naturally occurs, providing insights into real-life behaviors and interactions in their natural context.
Naturalistic observation is a research method commonly used by psychologists and other social scientists.
This technique involves observing and studying the spontaneous behavior of participants in natural surroundings. The researcher simply records what they see in whatever way they can.
In unstructured observations, the researcher records all relevant behavior with a coding system. Often, too much gets recorded. What is captured may not be the most important part, so the approach is usually used as a pilot study to see which behaviors are worth recording later.
Compared with controlled observations, it is like the difference between studying wild animals in a zoo and studying them in their natural habitat.
Human studies show the same pattern. Margaret Mead used this method to research the way of life of different tribes living on islands in the South Pacific. Kathy Sylva used it to study children at play by observing their behavior in a playgroup in Oxfordshire.
Collecting Naturalistic Behavioral Data
New technology enables unobtrusive naturalistic data collection.
The EAR is a small wearable recording device. It periodically samples ambient sounds to give a representative picture of daily life (Mehl et al., 2012).
EARs sample brief 30-50 second snippets several times an hour. Although coding the recordings requires extensive resources, EARs can capture spontaneous behaviors like arguments or laughter.
EARs minimize participant reactivity — the tendency for behavior to change simply because a person knows they are being observed — since sampling occurs outside of awareness. This reduces the Hawthorne effect.
The SenseCam is another wearable device that passively captures images documenting daily activities. Though primarily used in memory research currently (Smith et al., 2014), systematic sampling of environments and behaviors via the SenseCam could enable innovative psychological studies in the future.
Strengths
- Ecological Validity: Observing the natural flow of behavior in its own setting gives studies greater ecological validity.
- Generates Ideas: Like case studies, naturalistic observation often suggests new avenues of inquiry a researcher would not otherwise have thought of.
- Captures Real Behavior: Researchers can capture behaviors as they unfold in real time, including socially undesirable or complex behaviors people may not self-report accurately.
Limitations
- Small, Biased Samples: Naturalistic studies are often conducted on a small scale and may not represent wider society by age, gender, social class, or ethnicity.
- Lower Reliability: Because other variables cannot be controlled, another researcher may find it difficult to repeat the study in exactly the same way.
- Resource-Intensive: Training coders, maintaining inter-rater reliability, and preventing judgment drift are all time-consuming during the data coding phase.
- No Causation: Without manipulation of variables, cause-and-effect relationships cannot be established.
Participant Observation
Participant observation is a variant of the above (natural observations) but here, the researcher joins in and becomes part of the group they are studying to get a deeper insight into their lives.
If it were research on animals, we would now not only be studying them in their natural habitat but be living alongside them as well!
Perhaps the most famous — and most ethically contested — participant observation in psychology is David Rosenhan’s study of psychiatric admission.
- Aim: to test whether staff could spot “sane” from “insane.”
- Method: eight healthy volunteers, including Rosenhan himself, sought admission to twelve US psychiatric hospitals, each reporting a single fabricated symptom: hearing voices saying “empty,” “hollow,” and “thud.” Once admitted, they behaved entirely normally while secretly recording ward life.
- Results: every pseudopatient but one was admitted with a diagnosis of schizophrenia, and no staff member ever detected the deception. Ordinary behaviors, such as note-taking, were reinterpreted as symptoms.
- Conclusion: psychiatric diagnosis at the time could not reliably distinguish the sane from the insane once a diagnostic label had been applied (Rosenhan, 1973).
A second classic covert participant observation is Leon Festinger’s study of a small doomsday cult. Its members believed the world would end on a specific date, and that believers would be rescued by a spacecraft.
- Aim: to observe a group’s beliefs after a false prophecy fails.
- Method: Festinger and colleagues infiltrated the group as ordinary members, joining its meetings and preparations while secretly recording what they saw.
- Results: midnight passed with no apocalypse. Rather than abandon their belief, the leader announced that the group’s faith had saved the Earth, and the once-secretive cult began actively recruiting strangers.
- Conclusion: when a costly commitment is disconfirmed and cannot easily be undone, the resulting discomfort is usually resolved by reinterpreting the belief rather than abandoning it. This response was later formalized as cognitive dissonance (Festinger, Riecken, & Schachter, 1956).
Participant observations can be either covert or overt. Covert is where the study is carried out “undercover.” The researcher’s real identity and purpose are kept concealed from the group being studied.
The researcher takes a false identity and role, usually posing as a genuine member of the group.
On the other hand, overt is where the researcher reveals his or her true identity and purpose to the group and asks permission to observe.
Limitations
- Recording Difficulties: Researchers can’t take notes openly during covert observation without blowing their cover, so they must rely on memory, risking forgotten details and quotations.
- Loss of Objectivity: A researcher who becomes too involved risks selectively noticing what they expect or want to see rather than recording everything, which reduces validity.
Recording of Data
With controlled/structured observation studies, an important decision the researcher has to make is how to classify and record the data. Usually, this will involve a method of sampling.
In most coding systems, codes or ratings are made either per behavioral event or per specified time interval (Bakeman & Quera, 2011).
The three main sampling methods are:
- Event sampling. The observer decides in advance what types of behavior (events) she is interested in and records all occurrences. All other types of behavior are ignored.
Event-based coding involves identifying and segmenting interactions into meaningful events rather than timed units.
-
For example, parent-child interactions may be segmented into control or teaching events to code. Interval recording involves dividing interactions into fixed time intervals (e.g., 6-15 seconds) and coding behaviors within each interval (Bakeman & Quera, 2011).
-
Event recording allows counting event frequency and sequencing while also potentially capturing event duration through timed-event recording. This provides information on time spent on behaviors.
-
- Time (interval) sampling. The key feature of time sampling is that the interaction is divided into continuous fixed time intervals (e.g., every 15 seconds, 10 minutes every hour, 1 hour per day), and the observer codes for the behaviors that occur within each interval period.
- Interval recording is common in microanalytic coding to sample discrete behaviors in brief time samples across an interaction. The time unit can range from seconds to minutes to whole interactions. Interval recording requires segmenting interactions based on timing rather than events (Bakeman & Quera, 2011).
- Instantaneous (target time) sampling. The observer decides in advance the pre-selected moments when observation will occur and records what is happening at that instant. Everything happening before or after is ignored.
- Instantaneous sampling provides snapshot coding at certain moments rather than summarizing behavior within full intervals. This allows quicker coding but may miss behaviors in between target times.
Coding Systems
The coding system should focus on behaviors, patterns, individual characteristics, or relationship qualities that are relevant to the theory guiding the study (Wampler & Harper, 2014).
Codes vary in how much inference they require. A concrete, low-inference code might be frequency of eye contact; an abstract, high-inference code might be degree of rapport between a therapist and client (Hill & Lambert, 2004). More inference generally means lower reliability.
Coding schemes also vary in granularity. Micro-level schemes capture fine-grained behaviors, such as specific facial movements. Macro-level schemes code broader behavioral states or interactions instead.
The right level of detail depends on the research question and the study’s practical constraints.
Another consideration is how concrete the codes are. Some schemes use physically based codes that are directly observable, such as “eyes closed.” Others use socially based codes that require more inference, such as “showing empathy.”
This distinction matters. Physically based codes are easier to apply consistently, but socially based codes often capture more meaningful behavioral constructs.
Most coding schemes aim to be mutually exclusive and exhaustive (ME&E). This means only one code can apply at a time, and there is always an applicable code.
This property simplifies both coding and analysis.
For example, a simple ME&E set for coding infant state might include: 1) Quiet alert, 2) Crying, 3) Fussy, 4) REM sleep, and 5) Deep sleep. At any given moment, an infant would be in one and only one of these states.
Macroanalytic coding systems
Macroanalytic systems use large, broad coding units. They rate or summarize behavior patterns across long stretches of interaction, not small, discrete acts.
These systems focus on broad, overarching themes.
For example, a macroanalytic coding system may rate the overall degree of therapist warmth or client engagement for an entire therapy session. This requires coders to summarize and infer these constructs across the whole interaction, rather than coding smaller behavioral units. The trade-off is worth it for context.
These systems require observers to make more inferences, which is more time-consuming, but can better capture contextual factors, stability over time, and how behaviors interrelate (Carlson & Grotevant, 1987).
Examples of Macroanalytic Coding Systems:
- Emotional Availability Scales (EAS): This system assesses the quality of emotional connection between caregivers and children across dimensions like sensitivity, structuring, non-intrusiveness, and non-hostility.
- Classroom Assessment Scoring System (CLASS): Evaluates the quality of teacher-student interactions in classrooms across domains like emotional support, classroom organization, and instructional support.
Microanalytic coding systems
Microanalytic coding systems rate behaviors using small, discrete units.
These systems focus on capturing specific, discrete behaviors or events as they occur moment-to-moment. Coding often happens second-by-second.
For example, a microanalytic system may code each instance of eye contact or head nodding during a therapy session.
Microanalytic systems require less inference from coders and allow for analysis of behavioral contingencies and sequential interactions between therapist and client. However, they are more time-consuming and expensive to implement than macroanalytic approaches.
Examples of Microanalytic Coding Systems:
- Facial Action Coding System (FACS): Codes minute facial muscle movements to analyze emotional expressions.
- Specific Affect Coding System (SPAFF): Used in marital interaction research to code specific emotional behaviors.
- Noldus Observer XT: A software system that allows for detailed coding of behaviors in real-time or from video recordings.
Mesoanalytic coding systems
Mesoanalytic coding systems attempt to balance macro- and micro-analytic approaches.
In contrast to macroanalytic systems that summarize behaviors in larger chunks, mesoanalytic systems use medium-sized coding units that target more specific behaviors or interaction sequences (Bakeman & Quera, 2017).
For example, a mesoanalytic system may code each instance of a particular type of therapist statement or client emotional expression. However, mesoanalytic systems still use larger units than microanalytic approaches coding every speech onset/offset. This is a deliberate trade-off.
The goal of balancing specificity and feasibility makes mesoanalytic systems well-suited for many research questions (Morris et al., 2014). Mesoanalytic codes can preserve some sequential information while remaining efficient enough for studies with adequate but limited resources. Reliability matters here too.
For instance, a mesoanalytic couple interaction coding system could target key behavior patterns like validation sequences without coding turn-by-turn speech.
In this way, mesoanalytic coding allows reasonable reliability and specificity without requiring extensive training or observation. The mid-level focus offers a pragmatic compromise between depth and breadth in analyzing interactions.
Examples of Mesoanalytic Coding Systems:
- Feeding Scale for Mother-Infant Interaction: Assesses feeding interactions in 5-minute episodes, coding specific behaviors and overall qualities.
- Couples Interaction Rating System (CIRS): Codes specific behaviors and rates overall qualities in segments of couple interactions.
- Teaching Styles Rating Scale: Combines frequency counts of specific teacher behaviors with global ratings of teaching style in classroom segments.
Preventing Coder Drift
Coder drift results in a measurement error caused by gradual shifts in how observations get rated according to operational definitions, especially when behavioral codes are not clearly specified.
This type of error creeps in when coders fail to regularly review what precise observations constitute or do not constitute the behaviors being measured.
Preventing drift refers to taking active steps to maintain consistency and minimize changes or deviations in how coders rate or evaluate behaviors over time. Specifically, some key ways to prevent coder drift include:
- Operationalize codes: It is essential that code definitions unambiguously distinguish what interactions represent instances of each coded behavior.
- Ongoing training: Returning to those operational definitions through ongoing training serves to recalibrate coder interpretations and reinforce accurate recognition. Having regular “check-in” sessions where coders practice coding the same interactions allows monitoring that they continue applying codes reliably without gradual shifts in interpretation.
- Using reference videos: Coders periodically coding the same “gold standard” reference videos anchors their judgments and calibrate against original training. Without periodic anchoring to original specifications, coder decisions tend to drift from initial measurement reliability.
- Assessing inter-rater reliability: Statistical tracking that coders maintain high levels of agreement over the course of a study, not just at the start, flags any declines indicating drift. Sustaining inter-rater agreement requires mitigating this common tendency for observer judgment change during intensive, long-term coding tasks.
- Recalibrating through discussion: Having meetings for coders to discuss disagreements openly explores reasons judgment shifts may be occurring over time. Consensus on the application of codes is restored.
- Adjusting unclear codes: If reliability issues persist, revisiting and refining ambiguous code definitions or anchors can eliminate inconsistencies arising from coder confusion.
Essentially, the goal of preventing coder drift is maintaining standardization and minimizing unintentional biases that may slowly alter how observational data gets rated over periods of extensive coding.
Through the upkeep of skills, continuing calibration to benchmarks, and monitoring consistency, researchers can notice and correct for any creeping changes in coder decision-making over time.
Reducing Observer Bias
Observational research is prone to observer biases resulting from coders’ subjective perspectives shaping the interpretation of complex interactions (Burghardt et al., 2012). When coding, personal expectations may unconsciously influence judgments. However, rigorous methods exist to reduce such bias.
Coding Manual
Coding manuals reduce subjectivity. A detailed coding manual — a document defining exactly what each code means — clearly states what behaviors and interaction dynamics observers should code (Bakeman & Quera, 2011).
High-quality manuals have strong theoretical and empirical grounding, laying out explicit coding procedures and providing rich behavioral examples to anchor code definitions (Lindahl, 2001).
Clear delineation of the frequency, intensity, duration, and type of behaviors constituting each code facilitates reliable judgments and reduces ambiguity for coders. Without this clarity, raters apply codes inconsistently.
Coder Training
Competent coders require both interpersonal perceptiveness and scientific rigor (Wampler & Harper, 2014). Training thoroughly reviews the theoretical basis for coded constructs and teaches the coding system itself.
Multiple “gold standard” criterion videos demonstrate code ranges that trainees independently apply. Coders then meet weekly to establish reliability of 80% or higher agreement both among themselves and with master criterion coding (Hill & Lambert, 2004).
Ongoing training manages coder drift over time. Revisions to unclear codes may also improve reliability. Both careful selection and investment in rigorous training increase quality control.
Blind Methods
To prevent bias, coders should remain unaware of specific study predictions or participant details (Burghardt et al., 2012). Separate data gathering versus coding teams helps maintain blinding.
In addition, scheduling procedures can prevent coders from rating data collected directly from participants with whom they have had personal contact. Maintaining coder independence and blinding enhances objectivity.
Data Analysis Approaches
Data analysis in behavioral observation aims to transform raw observational data into quantifiable measures that can be statistically analyzed.
The choice of analysis approach is not arbitrary. It depends on the research questions, study design, and the nature of the data collected.
Different data types need different analytical approaches. Interval data records behavior at fixed time points; event data notes the occurrence of behaviors as they happen; and timed-event data captures both occurrence and duration.
The level of measurement matters too: categorical, ordinal, or continuous.
Researchers typically start with simple descriptive statistics to get a feel for their data before moving on to more complex analyses. This stepwise approach allows for a thorough understanding of the data and can often reveal unexpected patterns or relationships that merit further investigation.
simple descriptive statistics
Descriptive statistics give an overall picture of behavior patterns and are often the first step in analysis.
- Frequency counts tell us how often a particular behavior occurs, while rates express this frequency in relation to time (e.g., occurrences per minute).
- Duration measures how long behaviors last, offering insight into their persistence or intensity.
- Probability calculations indicate the likelihood of a behavior occurring under certain conditions, and relative frequency or duration statistics show the proportional occurrence of different behaviors within a session or across the study.
These simple statistics form the foundation of behavioral analysis, providing researchers with a broad picture of behavioral patterns.
They can reveal which behaviors are most common, how long they typically last, and how they might vary across different conditions or subjects.
Take a study of classroom behavior. These statistics might show how often students raise their hands, or how long they stay focused on a task. They might also show what proportion of time is spent on different activities.
contingency analyses
Contingency analyses help identify if certain behaviors tend to occur together or in sequence.
- Contingency tables, also known as cross-tabulations, display the co-occurrence of two or more behaviors, allowing researchers to see if certain behaviors tend to happen together.
- Odds ratios provide a measure of the strength of association between behaviors, indicating how much more likely one behavior is to occur in the presence of another.
- Adjusted residuals in these tables can reveal whether the observed co-occurrences are significantly different from what would be expected by chance.
In a study of parent-child interactions, contingency analyses might reveal whether a parent’s praise is more likely to follow a child’s successful completion of a task. They might equally reveal whether a child’s tantrum is more likely to occur after a parent refuses a request.
These analyses can uncover important patterns in social interactions, learning processes, or behavioral chains.
sequential analyses
Sequential analyses are crucial for understanding processes and temporal relationships between behaviors.
- Lag sequential analysis looks at the likelihood of one behavior following another within a specified number of events or time units.
- Time-window sequential analysis examines whether a target behavior occurs within a defined time frame after a given behavior.
These methods are particularly valuable for understanding processes that unfold over time, such as conversation patterns, problem-solving strategies, or the development of social skills.
observer agreement
Since human observers often code behaviors, it’s important to check reliability. This is typically done through measures of observer agreement.
- Cohen’s kappa is commonly used for categorical data, providing a measure of agreement between observers that accounts for chance agreement.
- Intraclass correlation coefficient (ICC): Used for continuous data or ratings.
Good observer agreement is crucial for the validity of the study, as it demonstrates that the observed behaviors are consistently identified and coded across different observers or time points.
advanced statistical approaches
As researchers delve deeper into their data, they often employ more advanced statistical techniques.
- Analysis of variance (ANOVA) can be used to compare behavior frequencies or durations across different groups or conditions.
- For instance, an ANOVA might reveal differences in the frequency of aggressive behaviors between children from different socioeconomic backgrounds or in different school settings.
- Multilevel modeling is particularly useful in behavioral observation studies where data is nested – for example, behaviors within individuals, individuals within groups, or observations across multiple time points.
- This approach allows researchers to account for dependencies in the data and to examine how behaviors might be influenced by factors at different levels (e.g., individual characteristics, group dynamics, and situational factors).
- Time series analysis is another powerful tool, especially for studies that involve continuous observation over extended periods.
- This method can reveal trends, cycles, or patterns in behavior over time, which might not be apparent from simpler analyses. For instance, in a study of animal behavior, time series analysis might uncover daily or seasonal patterns in feeding, mating, or territorial behaviors.
representation techniques
Representation techniques help organize and visualize data:
- Code-unit grid: Represents data as a matrix of behaviors and time units
- Many researchers use a code-unit grid, which represents the data as a matrix with behaviors as rows and time units as columns.
- This format facilitates many types of analyses and allows for easy visualization of behavioral patterns.
- Sequential Data Interchange Standard (SDIS): Standardizes data format for analysis
- Standardized formats like the Sequential Data Interchange Standard (SDIS) help ensure consistency in data representation across studies and facilitate the use of specialized analysis software.
- Indeed, the complexity of behavioral observation data often necessitates the use of specialized software tools. Programs like GSEQ, Observer, and INTERACT are designed specifically for the analysis of observational data and can perform many of the analyses described above efficiently and accurately.
Critical Evaluation of Observational Methods
Each type of observation above has its own strengths and limitations. Judged as a whole, though, the method faces five recurring challenges: ecological validity, representativeness, reliability, reactivity, and the ethics of covert study.
Ecological Validity
Observation’s central strength is ecological validity: because the researcher records behavior as it actually happens, findings often generalize better to everyday life than a tightly controlled experiment does.
This strength has to be earned, though, not assumed. A controlled observation carried out in an artificial laboratory setting, such as Bandura and colleagues’ Bobo doll study, sacrifices exactly this advantage for greater control over the situation.
This is not just theoretical. EAR-sampled conversation, laughter, and conflict predicted health outcomes as well as, or better than, what participants reported about themselves (Mehl et al., 2012).
That is a genuinely high bar: an unobtrusive recording device matching, or beating, a person’s own account of their day. It shows what real ecological validity can buy a researcher, when the design actually earns it.
Representativeness and Generalisability
Naturalistic observation is often conducted on a small scale, in a single setting. The resulting sample can be biased by age, gender, social class, or culture, which limits how confidently findings generalize to a wider population.
Cross-cultural research avoids this problem by deliberately sampling many settings rather than one convenient site. This is rare. It happens mainly because broader sampling costs far more to run.
A six-culture study of children’s prosocial behavior is a useful counter-example, since it deliberately built a broad, cross-cultural sample rather than settling for one convenient setting (Whiting & Whiting, 1975).
That kind of scale is rare for a reason. An early home-recording method, though innovative, stayed limited to a handful of families. The reason was simple: processing hours of recorded material took far too much time (Christensen, 1979).
Reliability of Coding
Reliability is not something observation gets for free. It depends entirely on the rigor of the coding process behind it.
High-inference, socially based codes, such as “showing empathy,” are harder to apply consistently than low-inference, physically based codes, such as “eyes closed.” Reliability must be actively maintained. It depends on ongoing training and inter-observer agreement checks, not assumption.
Burghardt and colleagues (2012) found that most published observational studies in their audited journals did not report blind coding or formal inter-observer reliability checks at all.
That gap is not a one-off oversight.
It points to reliability being a standing concern for the field as a whole, not a flaw confined to one weak study. The Contemporary Research evidence below tracks whether this has since improved.
Reactivity and Demand Characteristics
When participants know they are being watched, their behavior can change simply because of that knowledge. This is the Hawthorne effect, one example of the wider demand characteristics problem.
The effect is measurable. Children’s facial expressions of pain become more exaggerated in front of an audience, such as a parent during a medical exam (Vervoort et al., 2008). This is a sign that being watched changes what gets recorded.
Reactivity is not all-or-nothing. It is a matter of degree, tied to how overt the design is.
One study tested the EAR under a law that forced participants to wear a visible badge warning that their conversation might be recorded (Manson & Robbins, 2017). Even with this constant reminder, self-reported discomfort was no higher, and in some ways lower, than in the device’s original evaluation.
Covert Observation and Ethics
Covert observation trades reactivity for a serious ethical problem. Participants cannot consent to what they don’t know.
Two of psychology’s most striking participant-observation studies were covert precisely because overtness would have destroyed the phenomenon under study.
Staff could not have been tested on detecting a fake symptom if they had been warned to expect fakers. Equally, a doomsday cult could not have been observed reacting to a failed prophecy if it knew outsiders were watching (see the Rosenhan study above).
Neither study’s participants gave informed consent. That is why covert designs sit at the sharpest edge of the discipline’s ethics, even where they may be the only design able to answer the question.
Spitzer’s (1975) rebuttal of Rosenhan is a reminder that the ethical cost of deception is not automatically repaid by unambiguous scientific value.
Contemporary Research
Two strands of research since 2015 have tested the method’s own assumptions rather than simply asserting them.
Is Observer Bias Reporting Improving?
Burghardt and colleagues (2012) reviewed five major animal-behavior journals across five decades. Most published observational studies did not report two practices methodologists consider essential: coding blind to condition, and formally checking inter-observer reliability.
Freeberg, Benson, and Burghardt (2024) followed up directly. They recoded the same five journals’ 2020 volumes with identical criteria.
- Aim: to check whether reporting had improved since the 2012 audit.
- Method: the earlier content analysis was repeated exactly, with articles from the same five journals’ 2020 volumes coded using the same criteria, for a direct before-and-after comparison.
- Results: reporting rates had risen in all five journals, in some cases substantially, but still lagged behind standards in Infancy, a developmental journal using much of the same methodology.
- Conclusion: the field has genuinely improved but “has a long way to go” (Freeberg et al., 2024). Routine pre-registration and mandatory blind-coding reports are the next steps.
Naturalistic Observation of Family Life
This strand shows what naturalistic observation is used for now.
Wang and Repetti (2016) coded couples’ daily interactions. They found only partial support for the idea that partners give more support to a spouse who is visibly more stressed.
McNeil and Repetti (2021) used the same broad approach to catalogue which specific positive emotions actually occur in everyday family life, and how often.
Both questions are hard to answer with a questionnaire, since self-report is filtered through memory and social desirability. Naturalistic observation reaches behavior that self-report cannot.

Key Takeaways
- No Manipulation: Observation records behavior as it happens rather than manipulating an independent variable, so it cannot establish cause and effect the way an experiment can.
- Controlled vs Naturalistic: Controlled observation uses a standardized setting and schedule for easy comparison; naturalistic observation records behavior in its normal setting for higher ecological validity.
- Participant vs Non-Participant: A participant observer joins the group being studied for deeper insight, risking a loss of objectivity; a non-participant observer stays outside it.
- Covert vs Overt: Covert observation hides the researcher’s identity to reduce reactivity but raises serious consent issues; overt observation is transparent but risks the Hawthorne effect.
- Coding and Reliability: Behavior is recorded using a sampling method and a coding system, and reliability depends on rigorous observer training, not just a good schedule.
- Modern Evidence: Recent research shows observer-bias reporting is improving but still lags other fields, and unobtrusive tools like the EAR keep reactivity low even when disclosed.
References
Bakeman, R., & Quera, V. (2017). Sequential analysis and observational methods for the behavioral sciences. Cambridge University Press.
Burghardt, G. M., Bartmess-LeVasseur, J. N., Browning, S. A., Morrison, K. E., Stec, C. L., Zachau, C. E., & Freeberg, T. M. (2012). Minimizing observer bias in behavioral studies: A review and recommendations. Ethology, 118(6), 511-517.
Festinger, L., Riecken, H. W., & Schachter, S. (1956). When prophecy fails: A social and psychological study of a modern group that predicted the destruction of the world. University of Minnesota Press.
Freeberg, T. M., Benson, S. A., & Burghardt, G. M. (2024). Minimizing observer bias in animal behavior studies revisited: Improvement, but a long way to go. Ethology, 130(6). https://doi.org/10.1111/eth.13446
Hill, C. E., & Lambert, M. J. (2004). Methodological issues in studying psychotherapy processes and outcomes. In M. J. Lambert (Ed.), Bergin and Garfield’s handbook of psychotherapy and behavior change (5th ed., pp. 84–135). Wiley.
Lindahl, K. M. (2001). Methodological issues in family observational research. In P. K. Kerig & K. M. Lindahl (Eds.), Family observational coding systems: Resources for systemic research (pp. 23–32). Lawrence Erlbaum Associates.
Manson, J. H., & Robbins, M. L. (2017). New evaluation of the Electronically Activated Recorder (EAR): Obtrusiveness, compliance, and participant self-selection effects. Frontiers in Psychology, 8, Article 658. https://doi.org/10.3389/fpsyg.2017.00658
McNeil, G. D., & Repetti, R. L. (2021). Everyday emotions: Naturalistic observation of specific positive emotions in daily family life. Journal of Family Psychology, 35(2), 172–181. https://doi.org/10.1037/fam0000655
Mehl, M. R., Robbins, M. L., & Deters, F. G. (2012). Naturalistic observation of health-relevant social processes: The electronically activated recorder methodology in psychosomatics. Psychosomatic Medicine, 74(4), 410–417.
Morris, A. S., Robinson, L. R., & Eisenberg, N. (2014). Applying a multimethod perspective to the study of developmental psychology. In H. T. Reis & C. M. Judd (Eds.), Handbook of research methods in social and personality psychology (2nd ed., pp. 103–123). Cambridge University Press.
Rosenhan, D. L. (1973). On being sane in insane places. Science, 179(4070), 250–258. https://doi.org/10.1126/science.179.4070.250
Smith, J. A., Maxwell, S. D., & Johnson, G. (2014). The microstructure of everyday life: Analyzing the complex choreography of daily routines through the automatic capture and processing of wearable sensor data. In B. K. Wiederhold & G. Riva (Eds.), Annual Review of Cybertherapy and Telemedicine 2014: Positive Change with Technology (Vol. 199, pp. 62-64). IOS Press.
Spitzer, R. L. (1975). On pseudoscience in science, logic in remission, and psychiatric diagnosis: A critique of Rosenhan’s “On being sane in insane places”. Journal of Abnormal Psychology, 84(5), 442–452. https://doi.org/10.1037/h0077124
Traniello, J. F., & Bakker, T. C. (2015). The integrative study of behavioral interactions across the sciences. In T. K. Shackelford & R. D. Hansen (Eds.), The evolution of sexuality (pp. 119-147). Springer.
Vervoort, T., Goubert, L., Eccleston, C., Verhoeven, K., De Clercq, A., Buysse, A., & Crombez, G. (2008). The effects of parental presence upon the facial expression of pain: The moderating role of child pain catastrophizing. Pain, 138(2), 277–285. https://doi.org/10.1016/j.pain.2007.12.013
Wampler, K. S., & Harper, A. (2014). Observational methods in couple and family assessment. In H. T. Reis & C. M. Judd (Eds.), Handbook of research methods in social and personality psychology (2nd ed., pp. 490–502). Cambridge University Press.
Wang, S., & Repetti, R. L. (2016). Who gives to whom? Testing the support gap hypothesis with naturalistic observations of couple interactions. Journal of Family Psychology, 30(4), 492–502. https://doi.org/10.1037/fam0000196