Schedules of Reinforcement in Psychology (Examples)

Key Takeaways

  • Rule-based: a reinforcement schedule is a rule stating which instances of a behavior, if any, will be reinforced.
  • Two broad categories: reinforcement is either continuous (every response reinforced) or partial/intermittent (only some responses reinforced).
  • Four partial schedules: crossing fixed/variable with ratio/interval yields fixed-ratio, variable-ratio, fixed-interval and variable-interval schedules.
  • Each has a signature pattern: fixed-ratio shows a post-reinforcement pause, fixed-interval a “scallop,” variable-ratio a steady high rate, and variable-interval a steady low rate.
  • Variable-ratio is the most persistent: it is the most resistant schedule to extinction of all four, which helps explain the pull of gambling.

A schedule of reinforcement is a rule that states how often, and under what conditions, a reinforcer follows a behavior.

Ferster and Skinner’s landmark 1957 book, Schedules of Reinforcement, showed that different schedules produce distinct, predictable response patterns. How and when a behavior is reinforced shapes its strength and persistence, often more than the reward itself.

Introduction

A schedule of reinforcement is a component of operant conditioning (also known as instrumental conditioning). It consists of an arrangement to determine when to reinforce behavior, based on time or on the number of responses.

Operant Conditioning Quick facts

Schedules of reinforcement can be divided into two broad categories: continuous reinforcement, which reinforces a response every time, and partial reinforcement, which reinforces a response occasionally.

The type of reinforcement schedule used significantly impacts the response rate and resistance to the extinction of the behavior.

Research into schedules of reinforcement has yielded important implications for the field of behavioral science, including choice behavior, behavioral pharmacology, and behavioral economics.

Continuous Reinforcement

In continuous schedules, reinforcement is provided every single time after the desired behavior.

Due to the behavior being reinforced every time, the association is easy to make, and learning occurs quickly. However, this also means that extinction occurs quickly after reinforcement is no longer provided.

For Example

We can better understand the concept of continuous reinforcement by using candy machines as an example.

Candy machines are examples of continuous reinforcement because every time we put money in (behavior), we receive candy in return (positive reinforcement).

Japan Candy Machine

However, if a candy machine were to fail to provide candy twice in a row, we would likely stop trying to put money in (Myers, 2011).

We have come to expect our behavior to be reinforced every time it is performed and quickly grow discouraged if it is not.

Partial (Intermittent) Reinforcement Schedules

Unlike continuous schedules, partial schedules only reinforce the desired behavior occasionally rather than all the time. This leads to slower learning since it is initially more difficult to make the association between behavior and reinforcement.

However, partial schedules also produce behavior that is more resistant to extinction. Organisms are tempted to persist in their behavior in hopes that they will eventually be rewarded.

For instance, slot machines at casinos operate on partial schedules. They provide money (positive reinforcement) after an unpredictable number of plays (behavior). Hence, slot players are likely to continuously play slots in the hopes that they will gain money in the next round (Skinner, 1953).

Partial reinforcement schedules are the most common in everyday life. They vary along two dimensions.

The first is whether the number of responses required is fixed or variable. The second is whether reinforcement depends on the number of responses (ratio) or on the passage of time (interval).

Fixed Schedule

In a fixed schedule, the number of responses or amount of time between reinforcements is set and unchanging. The schedule is predictable.

Variable Schedule

In a variable schedule, the number of responses or amount of time between reinforcements changes randomly. The schedule is unpredictable.

Ratio Schedule

A ratio schedule reinforcement occurs after a certain number of responses have been emitted.

Interval Schedule

Interval schedules involve reinforcing a behavior after a period of time has passed.

Combinations of these four descriptors yield four kinds of partial reinforcement schedules: fixed-ratio, fixed-interval, variable-ratio and variable-interval.

Partial (Intermittent) Reinforcement Schedules

Fixed Interval Schedule

In operant conditioning, a fixed-interval schedule is when reinforcement is given to a response after a specific, predictable amount of time has passed.

Such a schedule results in a tendency for organisms to increase the frequency of responses closer to the anticipated time of reinforcement. However, immediately after being reinforced, the frequency of responses decreases.

The fluctuation in response rates means that a fixed-interval schedule will produce a scalloped pattern rather than steady rates of responding.

For Example

An example of a fixed-interval schedule would be a teacher giving students a weekly quiz every Monday.

Over the weekend, there is suddenly a flurry of studying for the quiz. On Monday, the students take the quiz and are reinforced for studying (positive reinforcement: receive a good grade; negative reinforcement: do not fail the quiz).

For the next few days, they relax after finishing the stressful experience, until the next quiz date draws near again.

Variable Interval Schedule

In operant conditioning, a variable-interval schedule is when reinforcement is provided for the first response after a random, unpredictable amount of time has passed.

This schedule produces a low, steady response rate since organisms are unaware of the next time they will receive reinforcers.

For Example

A pigeon in Skinner’s box has to peck a bar to receive a food pellet. It is given a food pellet after varying time intervals ranging from 2-5 minutes.

Skinner box or operant conditioning chamber experiment outline diagram. Labeled educational laboratory apparatus structure for mouse or rat experiment to understand animal behavior vector illustration

It is given a pellet after 3 minutes, then 5 minutes, then 2 minutes, etc. It will respond steadily since it does not know when its behavior will be reinforced.

Fixed Ratio Schedule

In operant conditioning, a fixed-ratio schedule reinforces behavior after a specified number of correct responses.

This kind of schedule results in high, steady rates of response. Organisms are persistent in responding because of the hope that the next response might be one needed to receive reinforcement.

For Example

An example of a fixed-ratio schedule is a dressmaker paid $500 after every 10 dresses they make.

After shipping a batch of ten and collecting the $500, they are likely to take a short break before starting the next batch.

Variable Ratio Schedule

A variable ratio schedule is a schedule of reinforcement where a behavior is reinforced after a random number of responses.

This type of schedule produces very consistent and high rates of responding, as the organism never knows exactly when reinforcement will occur. Due to its unpredictability, organisms persistently repeat the desired behavior, driven by the anticipation that the next response could lead to reinforcement.

For Example

An example of a variable-ratio schedule would be a child being given candy for every 3-10 pages of a book they read. For example, they are given candy after reading 5 pages, then 3 pages, then 7 pages, then 8 pages, etc.

The unpredictable reinforcement motivates them to keep reading, even if they are not immediately reinforced after reading one page.

Response Rates of Different Reinforcement Schedules

Ratio schedules, which are linked to the number of responses made, produce higher response rates than interval schedules.

As well, variable schedules produce more consistent behavior than fixed schedules; the unpredictability of reinforcement results in more consistent responses than predictable reinforcement (Myers, 2011).

Reinforcement Schedules Graph

Extinction of Responses Reinforced at Different Schedules

Resistance to extinction refers to how long a behavior continues to be displayed even after it is no longer being reinforced. A response high in resistance to extinction will take a longer time to become completely extinct.

Different schedules of reinforcement produce different levels of resistance to extinction. In general, schedules that reinforce unpredictably are more resistant to extinction.

Therefore, the variable-ratio schedule is more resistant to extinction than the fixed-ratio schedule. The variable-interval schedule is more resistant to extinction than the fixed-interval schedule as long as the average intervals are similar.

In the fixed-ratio schedule, resistance to extinction increases as the ratio increases. In the fixed-interval schedule, resistance to extinction increases as the interval lengthens in time.

Out of the four types of partial reinforcement schedules, the variable-ratio is the schedule most resistant to extinction. This can help to explain addiction to gambling.

Even as gamblers may not receive reinforcers after a high number of responses, they remain hopeful that they will be reinforced soon.

This persistence is the partial-reinforcement extinction effect (PREE): behavior reinforced only intermittently persists longer after reinforcement stops than behavior that was reinforced every time.

Aim: Humphreys (1939) tested why intermittently rewarded behavior resists extinction so strongly.

Method: Humphreys conditioned an eyelid-blink response, reinforcing it either every time (continuous) or through a random alternation of reinforced and non-reinforced trials (partial).

Results: The partially reinforced response extinguished far more slowly than the continuously reinforced one, an outcome so counterintuitive it became known as “Humphreys’ paradox.”

Conclusion: Continuous reinforcement makes a change in contingency easy to detect, so responding stops quickly once rewards end. Partial reinforcement already includes long unrewarded runs, so a similar run during extinction gives no clear signal that reward has ended for good.

Implications for Behavioral Psychology

In his article Schedules of Reinforcement at 50: A Retrospective Appreciation, Morgan (2010) reviews how schedules-of-reinforcement research has shaped behavioral science.

He highlights three areas where the framework remains productive: choice behavior, behavioral pharmacology and behavioral economics.

Choice Behavior

Behaviorists have long been interested in how organisms choose between alternatives and reinforcers. They have studied behavioral choice through the use of concurrent schedules.

Operating two schedules of reinforcement (often both variable-interval) at the same time lets researchers study how organisms allocate their behavior between the options.

Herrnstein (1961) discovered the matching law using concurrent schedules: an organism’s response rate to each option closely matches the proportion of reinforcement it has produced.

Herrnstein (1970) later generalized this into a quantitative version of the law of effect.

Baum (1974) then refined it with the generalized matching law. It adds two systematic deviations: bias, a constant preference for one option, and undermatching, responding less extremely than strict matching predicts.

For instance, Joe’s father almost always gives him money when he asks, but his mother rarely does.

Because asking his father is reinforced more often, Joe is more likely to ask him rather than his mother for money.

Individuals choose behavior that provides the largest reward. Their choices also depend on the reinforcement’s rate, quality, delay and required effort.

In short, people prefer rewards that are larger, better, sooner and easier to get.

Behavioral Pharmacology

Schedules of reinforcement are used to evaluate preference and abuse potential for drugs. One method used in behavioral pharmacological research to do so is through a progressive ratio schedule.

In a progressive-ratio schedule, the response requirement rises each time after reinforcement is attained. In pharmacology research, participants must make an increasing number of responses to earn an injection of a drug (reinforcement).

Under a progressive ratio schedule, a single injection may require up to thousands of responses. Participants are measured for the point where responding eventually stops, which is referred to as the “break point.”

Gathering data about the breakpoints of drugs allows for a categorization mirroring the abuse potential of different drugs. Using the progressive ratio schedule to evaluate drug preference and/or choice is now commonplace in behavioral pharmacology.

Behavioral Economics

Operant experiments offer an ideal way to study microeconomic behavior; participants can be viewed as consumers and reinforcers as commodities.

Through experimenting with different schedules of reinforcement, researchers can alter the availability or price of a commodity and track how response allocation changes as a result.

For example, researchers can change the ratio schedule, increasing or decreasing how many responses earn the reinforcer.

This studies elasticity: how much demand for the reinforcer changes as its “price” in effort changes.

Another example of the role reinforcement schedules play is in studying substitutability by making different commodities available at the same price (same schedule of reinforcement). By using the operant laboratory to study behavior, researchers have the benefit of being able to manipulate independent variables and measure the dependent variables.

Critical Evaluation

The schedules framework is one of psychology’s most precise and predictive bodies of law. But it inherits the limits of the behaviorist tradition it grew from.

Strengths

  • Precise and predictive: produces a distinctive, replicable pattern on the cumulative record for every schedule.
  • Wide real-world reach: explains gambling, pay structures, drug-taking and consumer demand, and underpins applied behavior-change programs.

Schedules give operant psychology genuine predictive power. Each one yields a distinctive, replicable signature on the cumulative record: the fixed-ratio break-and-run, the fixed-interval scallop, the steady variable-ratio rate (Ferster & Skinner, 1957).

Because these patterns recur across species and settings, the framework predicts the rate and persistence of behavior, not just that reward “works.” This precision is why schedules count among behaviorism’s most robust achievements.

The framework also reaches far beyond the laboratory. It explains gambling persistence, pay-structure effects, drug break points and consumer demand, and it underpins token economies and behavior-change programs (Ayllon & Michael, 1959; Morgan, 2010).

A growing body of research since 2015 extends this analysis to smartphones, social media and games. Their variable-reward notification and feed designs mirror slot-machine schedules and may foster compulsive use.

Limitations

  • Cognitive factors are neglected: humans can respond to a described rule rather than their actual reinforcement history.
  • Reductionist and ethically fraught: omits meaning and motivation, and hands institutions a tool for covert behavioral control.
  • Limited animal-to-human generalizability: the core findings come from rats and pigeons in tightly controlled chambers.

The “pure” schedule pattern often breaks down in people. Schedule effects come mainly from controlled animal chambers, but in humans, instructions, expectations and rule-governed behavior can override the contingencies.

A person who understands a schedule may respond to the rule itself rather than to their actual reinforcement history. This limits how directly schedule laws generalize to complex human behavior.

Explaining behavior solely through reinforcement schedules also omits meaning, motivation and cognition, reducing purposive human activity to external contingencies rather than choice.

The same power raises ethical concerns. Employers, institutions and technology companies can arrange schedules that shape behavior without a person’s awareness or consent. Deliberate variable-reward design in gambling and technology is a direct ethical charge against applied schedule engineering.

Schedule effects were also established chiefly with rats and pigeons in tightly controlled chambers. Human behavior is shaped by language, self-awareness and culture, so these “pure” patterns may not transfer directly to people.

Mini Quiz

Below are examples of schedules of reinforcement at work in the real world. Read the examples and then determine which kind of reinforcement schedule is being used.

References

Ayllon, T., & Michael, J. (1959). The psychiatric nurse as a behavioral engineer. Journal of the Experimental Analysis of Behavior, 2(4), 323–334.

Baum, W. M. (1974). On two types of deviation from the matching law: Bias and undermatching. Journal of the Experimental Analysis of Behavior, 22(1), 231–242.

Ferster, C. B., & Skinner, B. F. (1957). Schedules of reinforcement. Appleton-Century-Crofts.

Herrnstein, R. J. (1961). Relative and absolute strength of response as a function of frequency of reinforcement. Journal of the Experimental Analysis of Behavior, 4(3), 267–272.

Herrnstein, R. J. (1970). On the law of effect. Journal of the Experimental Analysis of Behavior, 13(2), 243–266.

Humphreys, L. G. (1939). The effect of random alternation of reinforcement on the acquisition and extinction of conditioned eyelid reactions. Journal of Experimental Psychology, 25(2), 141–158.

Morgan, D. L. (2010). Schedules of Reinforcement at 50: A Retrospective Appreciation. The Psychological Record, 60(1), 151–172.

Myers, David G. (2011). Psychology (10th ed.). Worth Publishers.

Skinner, B. F. (1953). Science and human behavior. Macmillan.

What are schedules of reinforcement?

Schedules of reinforcement are rules that control the timing and frequency of reinforcement delivery in operant conditioning. They include fixed-ratio, variable-ratio, fixed-interval, and variable-interval schedules, each dictating a different pattern of rewards in response to a behavior.

Which schedule of reinforcement is most resistant to the extinction of learned responses?

The variable-ratio schedule of reinforcement is the most resistant to extinction. This is because the reinforcement is given after an unpredictable number of responses, making it more difficult for the behavior to cease. Examples include gambling or lottery games, where a win is unpredictable but can occur anytime.

Saul McLeod, PhD

BSc (Hons) Psychology, MRes, PhD, University of Manchester

Chartered Psychologist (CPsychol)

Saul McLeod, PhD, is a qualified psychology teacher with over 18 years of experience in further and higher education. He has been published in peer-reviewed journals, including the Journal of Clinical Psychology.


Annabelle G.Y. Lim

Psychology Graduate

BA (Hons), Psychology, Harvard University

Annabelle G.Y. Lim is a graduate in psychology from Harvard University. She has served as a research assistant at the Harvard Adolescent Stress & Development Lab.