Piliavin, I. M., Rodin, J., & Piliavin, J. A. (1969). Good Samaritanism: An underground phenomenon? Journal of Personality and Social Psychology, 13(4), 289–299.

The Piliavin, Rodin, and Piliavin (1969) subway Samaritan study is a covert field experiment on the New York underground that tested when bystanders help a stranger who collapses in public.
Passengers helped an apparently ill victim on 62 of 65 trials, far more often and faster than a drunk victim, and larger crowds did not reduce helping. The findings challenged the idea that groups always diffuse responsibility, and led the researchers to propose the arousal:cost–reward model of helping.
Aim
This study investigated how bystanders react when a stranger collapses on a train.
The researchers tested four factors thought to influence helping:
- Would an ill person get more help than a drunk person (the type of victim)?
- Would people help others of the same race before helping those of different races?
- If a model person started helping the victim, would that encourage others to also help?
- Would the number of bystanders who saw the victim influence how much help was given?
Procedure
This study was a field experiment on a 7 ½ minute non-stop journey on a New York underground train, using various coaches along the train. Participants were passengers who were on board.
A team of 4 university students (a male victim, a male model, and 2 female observers) staged the emergency to observe how passengers would react.
A ‘victim’ staged an ‘emergency’ by collapsing (in the designated ‘critical area’).

After collapsing, the victim lay on his back on the floor. If not helped earlier in the journey by a participant or model, the model assisted the victim at the end of the journey.
Participants’ reactions were then watched by covert observers.
Participants
The victims were: males aged 26 -35; three white, one black; identically dressed in a US army-style jacket, old trousers, and no tie.
The ‘drunk’ smelled of alcohol and carried a spirits bottle wrapped in a brown paper bag (38 trials). The ‘ill’ victim appeared sober and carried a black cane (65 trials).
The models were males aged 24 – 29, wore casual but not identical clothes, and helped by raising the victim to a sitting position and staying with him.
Independent Variables
- Type of victim (drunk or appearing ill and would hold a walking cane).
- Race of victim (black or white).
- Effect of the model (a confederate in disguise): helped early (about 70 seconds in) or late (about 150 seconds in), from near the victim (‘critical area’) or further away (‘adjacent’), or was absent altogether. The model condition was assigned randomly.
- Size of the witnessing group (a naturally occurring independent variable). The sample consisted of the 4450 American passengers using that particular train, 45% of which were black and 55% white.
Dependent Variables
The dependent variables were covertly recorded (from behind newspapers) by two female observers seated in the adjacent area (during 103 victim trials):
- Number of bystanders
- Frequency of witness help
- Latency (time) to help
- Race of helper
- Sex of helper
- Number of helpers
- Movement out of the critical area
- Any verbal comments made by bystanders
There were 6-8 trials per day, on journeys in alternating directions, all the same victim type on any day.
Findings
- Models were rarely needed; the public usually helped quickly on their own.
- Apparently ill victims were helped much more often, and much faster, than apparently drunk victims: 62 of 65 trials for the ill victim compared with 19 of 38 for the drunk victim, with a median help latency of about 5 seconds for the ill victim against roughly 109 seconds for the drunk victim.
- Males are more likely to help than females (60% of travelers were male, but 90% of first helpers were male).
- Race has little effect on helping, although a drunk victim is less likely to receive opposite-race help.
- The longer no help is offered, the less important modeling becomes and the more likely someone is to leave the area, and more so with drunk victims.
- Spontaneous comments were more common in the drunk condition.
Conclusion
One of the surprising findings in this study was that there was no diffusion of responsibility. The size of the group made no difference in how much help a victim received. Piliavin et al. offered several explanations for this:
- Passengers were trapped on the train and could not really leave the situation; on the street, the results might have been different.
- Helping cost passengers little effort, since they were already seated and waiting for the next stop anyway.
- Unlike the Kitty Genovese case, it was clear what the problem was for the bystanders sitting next to the victim.
The Arousal:Cost–Reward Model
Piliavin et al. (1969) proposed the arousal:cost–reward model. It is a major alternative to Latané and Darley’s decision model. They called it a ‘fine tuning’ of that earlier account. Like the decision model, it has two stages before a bystander acts.
Stage 1: Emotional Arousal
Seeing someone in distress produces an unpleasant emotional response. This physiological arousal is the model’s basic motivational engine.
Without arousal, a bystander has little reason to act. The greater the arousal, the more likely a bystander is to help.
Acting is simply the fastest way to switch the discomfort off. Several factors intensify this arousal.
Empathy with the victim raises it further. Physical closeness to the emergency matters too. So does how long the crisis drags on.
This first stage supplies only the motivation to act. A separate, cognitive stage decides exactly what a bystander does next.
Without that second stage, arousal alone cannot explain who helps, or how.
Arousal, in short, is the engine. It is not yet the steering wheel.
Stage 2: The Cost-Reward Calculation
Arousal supplies the motivation to act. A second, cognitive stage decides exactly what to do.
The bystander weighs the costs and rewards of each option. This often happens fast, and not fully consciously.
They pick whichever option cuts the arousal fastest, for the fewest costs.
- Costs of helping: effort, time, lost resources, risk of harm, and unpleasant emotions such as disgust or embarrassment.
- Rewards of helping: gratitude from the victim and onlookers, social approval, and the self-satisfaction of having acted.
- Costs of not helping: guilt, the disapproval of others, and damage to one’s self-esteem.
These weights differ between people. So the model predicts patterns, not certainties.
Help was more likely for the ill victim than the drunk one. The drunk victim raised the cost, and lowered the sympathy, of helping.
Men offered help more often too. Their social role reduces self-blame, even though their perceived risk is higher.
Same-race bystanders were also likelier to help a drunk victim. Perceived risk is higher then, and same-race trust may matter.
A Worked Example
Picture a commuter watching the cane-carrying, sober victim collapse. The sight produces an immediate jolt of arousal.
The commuter is trapped in the carriage. Walking away will not switch the feeling off.
The victim is clearly sober and ill. Helping costs little: modest effort, little personal risk.
Not helping costs more: guilt, and disapproving looks in an enclosed space.
The cheapest way to switch off the arousal is to step forward. That is exactly what passengers overwhelmingly did.
Now replay the scene with the drunk victim instead. The smell of alcohol and the bottle raise the cost of helping.
Disgust rises. So does the risk of an unpredictable reaction.
The cost of not helping falls too. The victim seems partly responsible for his own state. This lowers sympathy for him.
Cheaper routes to reducing arousal now look attractive: looking away, or reinterpreting him as “just a drunk”.
Help becomes slower and rarer as a result.
Later Development of the Model
Piliavin et al.’s original two-stage sketch was elaborated over the following decade.
In a further subway experiment, Piliavin, Piliavin, and Rodin (1975) gave the victim a disfiguring facial birthmark. This raised the perceived cost of approaching him.
Help fell to about 61% of trials.
Diffusion of responsibility, which the 1969 study had failed to detect, began to reappear too. It returned once the costs of helping rose.
The model was elaborated most fully in the book Emergency Intervention (Piliavin, Dovidio, Gaertner, & Clark, 1981).
It added many interacting factors: the bystander’s own state, and the victim’s characteristics.
Together, these shape both how aroused a bystander becomes, and how they weigh the costs and rewards of helping.
Dovidio et al. (1991) reviewed two decades of evidence. They reached a clear conclusion.
Emotional reactions to another person’s distress, they found, do play an important role in motivating helping.
This finding supports the model’s core logic.
Critical Evaluation
Piliavin et al.’s study has real strengths and real limitations. Later research has also tested how well its central finding has held up. These are summarised below.
- Ecological Validity: a genuine field emergency, covertly observed, means the high helping rates reflect how people really behave, not how they think they should behave in a lab.
- Reliability and Rich Data: two independent observers recording about 4,450 passengers across 103 trials gave both statistical and qualitative data, and let the researchers check inter-rater reliability.
- Ethical Issues: passengers gave no informed consent, could not withdraw from a staged collapse, and could never be debriefed afterwards.
- Loss of Control and Generalisability: carriage composition, time of day, and passenger mood could not be controlled, and the sample was drawn from one city, line, and era.
- The Confound of Victim Characteristics: the drunk and ill victims differed on more than sobriety alone, making it hard to isolate which single feature reduced helping.
- An “Overly Calculating” Model?: critics argue the model makes helping sound too rational, and that arousal and helping are only correlated, not proven to cause one another.
- Contemporary Research Supports the Finding: CCTV analysis of real public disputes confirms that bystanders usually intervene, even in large crowds, echoing Piliavin et al.’s result.
Ecological Validity
The study’s biggest strength is its ecological validity. Real subway passengers faced what looked like a genuine collapse.
They were not volunteers who had signed up for an experiment.
The observers recorded behaviour covertly, from behind newspapers. None of the passengers knew they were being watched.
So there were no demand characteristics steering their behaviour.
This is why the very high helping rates, 62 of 65 trials for the ill victim, likely reflect real emergency behaviour. They are not rehearsed lab politeness.
The trade-off is a loss of control. Carriage composition, time of day, and passengers’ moods varied freely.
None of that could be held constant. A field experiment like this one buys realism at some cost to the tightness of its causal claims.
Reliability and Rich Data
About 4,450 passengers took part in the study.
That is a large sample, for a field experiment.
It generated two kinds of evidence: qualitative and quantitative.
The quantitative data, whether and how quickly someone helped, let the researchers run statistical comparisons across victim type, race, and group size.
The qualitative data came from spontaneous comments passengers made as the scene unfolded.
Together they build a fuller picture.
These add insight into why people acted, or held back. A frequency count alone cannot show that.
Two independent observers recorded each trial, not one. This let the researchers check inter-rater reliability.
That means how consistent different observers are when recording the same event.
Scale, two data types, and a reliability check together make this an unusually well-evidenced field study.
Ethical Issues
This raises some of psychology’s clearest ethical issues.
None of the passengers gave informed consent. They boarded expecting an ordinary commute.
They had no idea that an emergency, and their own reactions to it, were being staged and recorded.
They also could not withdraw once the ‘victim’ collapsed. Walking away from a moving train is not realistic.
Most seriously, they could never be debriefed. Trials ran continuously on anonymous commuters.
Passengers dispersed at the next stop, so there was no way to find them again afterwards.
Some may have carried genuine distress, or lingering guilt over not helping.
There was no chance for the researchers to check they were unharmed.
Against this, the study caused no physical harm. It produced knowledge about real helping that a consenting, debriefed lab sample could not have provided.
Loss of Control and Generalisability
As a field experiment, the study sacrificed control over many factors. These vary from trial to trial: carriage crowding, time of day, passengers’ moods.
None of these extraneous variables could be held constant.
So it is hard to be certain that victim type and crowd size, rather than a background factor, drove any single trial’s outcome.
The sample, though large at around 4,450 passengers, came from one city, one train line, and one culture.
All of that at a single point in time.
New York subway commuters of that era are not necessarily representative of bystanders elsewhere.
The researchers themselves noted that results on an open street, where people can simply walk away, might well differ.
Helping norms are also shaped by culture and history. The study’s figures should not be read as universal constants.
The Confound of Victim Characteristics
The victims did not differ cleanly.
The drunk victim smelled of alcohol and carried a bottle in a paper bag. He also invited a moral judgement that the sober, cane-carrying victim did not.
Several features changed at once.
Sobriety, smell, the prop carried, and the blame a bystander assigns all shifted together.
So it is hard to isolate which single feature reduced helping.
Disgust, risk, or blame could each be responsible.
A further limitation: the same small pool of student confederates played every victim across all 103 trials.
Actors differ too.
Individual differences between them could also have shaped how passengers responded.
Realism is gained in a field study like this one. But something is lost too.
The ability to isolate exactly which variable is doing the causal work.
An “Overly Calculating” Model?
Critics argue the arousal:cost–reward model makes bystanders sound too calculating.
In a real emergency, people do not consciously tot up the pros and cons of helping this deliberately.
There is also a causal-inference problem. Arousal and helping are typically only correlated in this kind of field data.
Yet the model treats arousal as the actual cause of helping.
Dovidio et al. (1991) reviewed two decades of follow-up evidence. They defended the model.
Emotional reactions to another person’s distress, they concluded, do play a genuine role in motivating people to help.
Bystanders, on this view, choose whichever response cuts their arousal fastest and at the least cost.
The weighing happens quickly, and largely without conscious awareness.
So the “accountancy” language is a loose metaphor. It stands for a real, fast psychological process, not a literal ledger.
Contemporary Research Supports the Finding
The clearest modern test of Piliavin et al.’s central claim comes from Philpot et al. (2020).
They analysed 219 real public disputes captured on CCTV. The footage came from the UK, the Netherlands, and South Africa.
At least one bystander intervened in about 90% of incidents.
The average event drew nearly four interveners.
Echoing the 1969 subway result, larger crowds were, if anything, more likely to produce a helper, not less.
This evidence is drawn from thousands of real bystanders, not laboratory volunteers or a single subway line.
The parallel to 1969 is striking.
It strongly supports Piliavin et al.’s original, surprising finding: intervention, not apathy, is the norm in a clear, unavoidable emergency.
Half a century on, the subway study’s most striking result still holds up against real-world data.
Groups help. They do not simply freeze.
Key Takeaways
- Field Experiment: Piliavin, Rodin, and Piliavin (1969) staged a covert emergency, a man collapsing, on the New York subway to see how real passengers would react.
- High Helping Rates: passengers helped the apparently ill victim on 62 of 65 trials, and models were rarely needed because someone usually stepped in first.
- Little Diffusion of Responsibility: crowd size made almost no difference to whether or how fast help arrived, contradicting earlier laboratory bystander studies.
- Ill vs Drunk: an apparently drunk victim was helped far less often (19 of 38 trials) and much more slowly than the ill victim.
- Arousal:Cost–Reward Model: the study introduced a two-stage model in which physiological arousal motivates action and a rapid cost-reward calculation decides how to respond.
- Modern Evidence: CCTV analysis of real public disputes (Philpot et al., 2020) confirms that bystander intervention, not apathy, is the norm even in large crowds.
References
Dovidio, J. F., Piliavin, J. A., Gaertner, S. L., Schroeder, D. A., & Clark, R. D. (1991). The arousal: Cost–reward model and the process of intervention: A review of the evidence. In M. S. Clark (Ed.), Prosocial behavior (pp. 86–118). Sage.
Piliavin, I. M., Piliavin, J. A., & Rodin, J. (1975). Costs, diffusion, and the stigmatized victim. Journal of Personality and Social Psychology, 32(3), 429–438. https://doi.org/10.1037/h0077092
Piliavin, I. M., Rodin, J., & Piliavin, J. A. (1969). Good Samaritanism: An underground phenomenon? Journal of Personality and Social Psychology, 13(4), 289–299.
Piliavin, J. A., Dovidio, J. F., Gaertner, S. L., & Clark, R. D. (1981). Emergency intervention. Academic Press.