Schedules of Reinforcement in Psychology (Examples)

A schedule of reinforcement is a rule for exactly how and when a reinforcer follows a behavior. Different schedules create very different patterns of responding, and some make a habit far harder to break than others.

Key Takeaways

  • Rule-based: a reinforcement schedule is a rule stating which instances of a behavior, if any, will be reinforced.
  • Two broad categories: reinforcement is either continuous (every response reinforced) or partial/intermittent (only some responses reinforced).
  • Four partial schedules: crossing fixed/variable with ratio/interval yields fixed-ratio, variable-ratio, fixed-interval and variable-interval schedules.
  • Each has a signature pattern: fixed-ratio shows a post-reinforcement pause, fixed-interval a “scallop,” variable-ratio a steady high rate, and variable-interval a steady low rate.
  • Variable-ratio is the most persistent: it is the most resistant schedule to extinction of all four, which helps explain the pull of gambling.

A schedule of reinforcement is a rule that states how often, and under what conditions, a reinforcer follows a behavior.

Ferster and Skinner’s landmark 1957 book, Schedules of Reinforcement, mapped out how these schedules operate within operant conditioning, showing that each produces a distinct, predictable response pattern. How and when a behavior is reinforced shapes its strength and persistence, often more than the reward itself.

Introduction

A schedule of reinforcement is a component of operant conditioning (also known as instrumental conditioning). It consists of an arrangement to determine when to reinforce behavior, based on time or on the number of responses.

Operant Conditioning Quick facts

Schedules of reinforcement can be divided into two broad categories: continuous reinforcement, which reinforces a response every time, and partial reinforcement, which reinforces a response occasionally.

The type of reinforcement schedule used significantly impacts the response rate and resistance to the extinction of the behavior.

Research into schedules of reinforcement has yielded important implications for the field of behavioral science, including choice behavior, behavioral pharmacology, and behavioral economics.

Continuous Reinforcement

In continuous schedules, reinforcement is provided every single time after the desired behavior.

Due to the behavior being reinforced every time, the association is easy to make, and learning occurs quickly. However, this also means that extinction occurs quickly after reinforcement is no longer provided.

For Example

We can better understand the concept of continuous reinforcement by using candy machines as an example.

Candy machines are examples of continuous reinforcement because every time we put money in (behavior), we receive candy in return (positive reinforcement).

Japan Candy Machine

However, if a candy machine were to fail to provide candy twice in a row, we would likely stop trying to put money in (Myers, 2011).

We have come to expect our behavior to be reinforced every time it is performed and quickly grow discouraged if it is not.

Partial (Intermittent) Reinforcement Schedules

Unlike continuous schedules, partial schedules only reinforce the desired behavior occasionally rather than all the time. This leads to slower learning since it is initially more difficult to make the association between behavior and reinforcement.

However, partial schedules also produce behavior that is more resistant to extinction. Organisms are tempted to persist in their behavior in hopes that they will eventually be rewarded.

For instance, slot machines at casinos operate on partial schedules. They provide money (positive reinforcement) after an unpredictable number of plays (behavior). Hence, slot players are likely to continuously play slots in the hopes that they will gain money in the next round (Skinner, 1953).

Partial reinforcement schedules are the most common in everyday life. They vary along two dimensions.

The first is whether the number of responses required is fixed or variable. The second is whether reinforcement depends on the number of responses (ratio) or on the passage of time (interval).

Fixed Schedule

In a fixed schedule, the number of responses or amount of time between reinforcements is set and unchanging. The schedule is predictable.

Variable Schedule

In a variable schedule, the number of responses or amount of time between reinforcements changes randomly. The schedule is unpredictable.

Ratio Schedule

A ratio schedule reinforcement occurs after a certain number of responses have been emitted.

Interval Schedule

Interval schedules involve reinforcing a behavior after a period of time has passed.

Combinations of these four descriptors yield four kinds of partial reinforcement schedules: fixed-ratio, fixed-interval, variable-ratio and variable-interval.

Partial (Intermittent) Reinforcement Schedules

Fixed Interval Schedule

In operant conditioning, a fixed-interval schedule is when reinforcement is given to a response after a specific, predictable amount of time has passed.

Such a schedule results in a tendency for organisms to increase the frequency of responses closer to the anticipated time of reinforcement. However, immediately after being reinforced, the frequency of responses decreases.

The fluctuation in response rates means that a fixed-interval schedule will produce a scalloped pattern rather than steady rates of responding.

For Example

An example of a fixed-interval schedule would be a teacher giving students a weekly quiz every Monday.

Over the weekend, there is suddenly a flurry of studying for the quiz. On Monday, the students take the quiz and are reinforced for studying (positive reinforcement: receive a good grade; negative reinforcement: do not fail the quiz).

For the next few days, they relax after finishing the stressful experience, until the next quiz date draws near again.

Variable Interval Schedule

In operant conditioning, a variable-interval schedule is when reinforcement is provided for the first response after a random, unpredictable amount of time has passed.

This schedule produces a low, steady response rate since organisms are unaware of the next time they will receive reinforcers.

For Example

A pigeon in Skinner’s box has to peck a bar to receive a food pellet. It is given a food pellet after varying time intervals ranging from 2-5 minutes.

Skinner box or operant conditioning chamber experiment outline diagram. Labeled educational laboratory apparatus structure for mouse or rat experiment to understand animal behavior vector illustration

It is given a pellet after 3 minutes, then 5 minutes, then 2 minutes, etc. It will respond steadily since it does not know when its behavior will be reinforced.

People show the same pattern with their phones. Checking email or social media for messages that arrive at unpredictable times produces the same low, steady rate of checking throughout the day.

Fishing works the same way; a bite can come after any unpredictable wait. A teacher’s surprise pop quizzes keep students studying steadily instead of cramming.

Fixed Ratio Schedule

In operant conditioning, a fixed-ratio schedule reinforces behavior after a specified number of correct responses.

This kind of schedule results in high, steady rates of response. Organisms are persistent in responding because of the hope that the next response might be one needed to receive reinforcement.

For Example

An example of a fixed-ratio schedule is a dressmaker paid $500 after every 10 dresses they make.

After shipping a batch of ten and collecting the $500, they are likely to take a short break before starting the next batch. This post-reinforcement pause is the schedule’s signature.

The same logic drives ordinary factory piece-rates and sales commission paid per fixed number of units. A teacher who awards a spelling star for every five correctly spelled words is using the identical rule. The more a person produces, the more reinforcement they earn, so output stays high.

Variable Ratio Schedule

A variable ratio schedule is a schedule of reinforcement where a behavior is reinforced after a random number of responses.

This type of schedule produces very consistent and high rates of responding, as the organism never knows exactly when reinforcement will occur. Due to its unpredictability, organisms persistently repeat the desired behavior, driven by the anticipation that the next response could lead to reinforcement.

For Example

An example of a variable-ratio schedule is a child given candy for reading pages of a book.

They might get candy after 5 pages, then after 3, then 7, then 8. The reward comes at no fixed point.

This unpredictable reinforcement keeps them reading, even on pages that bring no reward.

Compound Schedules

Beyond the four simple schedules, researchers have combined the basic rules into several compound schedules. Three matter most.

Each extends Ferster and Skinner’s (1957) original framework.

Differential Reinforcement (DRL, DRH and DRO)

Differential-reinforcement schedules reward a response based on how fast or how often it occurs, not just on whether it occurs at all.

DRL (differential reinforcement of low rates) reinforces a response only once a minimum time has passed since the last one. This teaches an organism to respond slowly and deliberately.

DRH (differential reinforcement of high rates) does the opposite: it reinforces only bursts of rapid responding.

A related schedule, DRO (differential reinforcement of other behavior), reinforces the absence of a target response.

It reduces problem behavior without punishment. A teacher might use it to reinforce a student for periods without talking out of turn.

Together, DRL, DRH and DRO show that schedules can shape not just whether a behavior happens, but its pace.

Interlocking, Multiple and Chained Schedules

Other schedules combine two component schedules instead.

In an interlocking schedule, the ratio and interval requirements are made interdependent, so meeting one requirement changes the other.

Multiple and mixed schedules alternate between two component schedules. A multiple schedule signals the switch with a discriminative stimulus; a mixed schedule gives no such signal.

This distinguishes discriminated responding from purely consequence-driven behavior.

A chained schedule requires an organism to complete a sequence of component schedules before it reaches a single terminal reinforcer. These arrangements let researchers study how stimuli, sequences and past reinforcement history combine to control behavior.

A chained schedule shows up in everyday routines too. Getting dressed, making coffee and leaving for work might all have to happen in sequence before a single paycheck follows.

Non-Contingent Reinforcement: Skinner’s “Superstition” Experiment

Not every schedule depends on behavior at all.

In a fixed-time or variable-time schedule, the reinforcer is delivered on a purely temporal basis, regardless of whether the organism responds. There is no genuine link between the response and the reward.

Skinner (1948) showed this vividly. He ran a well-known “superstition” experiment: hungry pigeons received food at a fixed interval, about every 15 seconds, no matter what they were doing.

The pigeons developed their own rituals. They turned counter-clockwise, tossed their heads, and swayed their bodies like a pendulum.

Whatever a bird happened to be doing when food arrived was accidentally, or adventitiously, reinforced. The finding shows that a reinforcement schedule can shape and sustain behavior even when the behavior itself has no real effect on when the reward arrives.

Together, these compound schedules show real flexibility in the model. Researchers can combine, chain, or strip away the response-reinforcer link entirely, and behavior still organizes itself around whatever schedule is in play.

Response Rates of Different Reinforcement Schedules

Ratio schedules, which are linked to the number of responses made, produce higher response rates than interval schedules.

As well, variable schedules produce more consistent behavior than fixed schedules; the unpredictability of reinforcement results in more consistent responses than predictable reinforcement (Myers, 2011).

Reinforcement Schedules Graph

Extinction of Responses Reinforced at Different Schedules

Resistance to Extinction Across Schedules

Resistance to extinction is how long a behavior continues once reinforcement stops. A response high in resistance takes longer to disappear.

Different schedules build different levels of resistance. In general, schedules that reinforce unpredictably resist extinction more than predictable ones do.

The variable-ratio schedule resists extinction more than the fixed-ratio schedule. The variable-interval schedule resists extinction more than the fixed-interval schedule, as long as the average intervals are similar.

Within a fixed schedule, resistance grows as the requirement grows: a larger fixed ratio, or a longer fixed interval, both increase resistance to extinction.

Variable-ratio is the most resistant of all four. This helps explain the pull of gambling addiction. A slot machine (variable-ratio) is far harder to walk away from than a vending machine (fixed-ratio) for exactly this reason.

The Partial-Reinforcement Extinction Effect

Gamblers often go through long losing streaks with no pay-off. They keep playing anyway.

They stay hopeful that reinforcement will come soon. This persistence is the partial-reinforcement extinction effect (PREE): behavior reinforced only intermittently persists longer after reinforcement stops than behavior that was reinforced every time.

One classic study shows why.

Aim: Humphreys (1939) tested why intermittently rewarded behavior resists extinction so strongly.

Method: Humphreys conditioned an eyelid-blink response, reinforcing it either every time (continuous) or through a random alternation of reinforced and non-reinforced trials (partial).

Results: The result was striking. The partially reinforced response extinguished far more slowly than the continuously reinforced one. This was so counterintuitive it became known as “Humphreys’ paradox.”

Conclusion: Continuous reinforcement makes a change in contingency easy to detect, so responding stops quickly once rewards end. Partial reinforcement already includes long unrewarded runs, so a similar run during extinction gives no clear signal that reward has ended for good.

Implications for Behavioral Psychology

In his article Schedules of Reinforcement at 50: A Retrospective Appreciation, Morgan (2010) reviews how schedules-of-reinforcement research has shaped behavioral science.

He highlights three areas where the framework remains productive: choice behavior, behavioral pharmacology and behavioral economics.

Choice Behavior

Behaviorists have long been interested in how organisms choose between alternatives and reinforcers. They study this choice using concurrent schedules.

Two or more schedules run at the same time on different response options. Operating two schedules at once, often both variable-interval, lets researchers see how organisms allocate their behavior between the options. This is choice behavior.

Herrnstein (1961) discovered the matching law using concurrent schedules: an organism’s response rate to each option closely matches the proportion of reinforcement it has produced.

Herrnstein (1970) later generalized this into a quantitative version of the law of effect.

Baum (1974) refined it further. His generalized matching law adds two systematic deviations: bias, a constant preference for one option, and undermatching, responding less extremely than strict matching predicts.

Take Joe, for instance. His father almost always gives him money when he asks, but his mother rarely does.

Because asking his father is reinforced more often, Joe is more likely to ask him rather than his mother for money.

Individuals choose behavior that provides the largest reward. Their choices also depend on the reinforcement’s rate, quality, delay and required effort.

In short, people prefer rewards that are larger, better, sooner and easier to get.

Behavioral Pharmacology

Schedules of reinforcement are used to evaluate preference and abuse potential for drugs. One key tool is the progressive-ratio schedule.

In a progressive-ratio schedule, the response requirement rises each time after reinforcement is attained. In pharmacology research, participants must make an increasing number of responses to earn an injection of a drug (reinforcement).

A single injection may eventually need thousands of responses. Participants are measured for the point where responding stops, called the “break point.”

Gathering data about drugs’ breakpoints allows researchers to categorize their abuse potential. Using the progressive-ratio schedule to assess drug preference is now standard practice in behavioral pharmacology.

Behavioral Economics

Operant experiments are an ideal way to study microeconomic behavior. Participants act as consumers, and reinforcers as commodities.

Researchers alter a commodity’s availability or price by changing the reinforcement schedule, then track how response allocation shifts.

They can change the ratio schedule, for example, so more or fewer responses earn the reinforcer. This measures elasticity: how much demand changes as effort “price” rises or falls.

Price matters here. Another example is substitutability, tested by offering different commodities at the same response price.

The operant laboratory lets researchers manipulate independent variables directly and measure the dependent variables that follow.

Critical Evaluation

The schedules framework is one of psychology’s most precise and predictive bodies of law. But it inherits the limits of the behaviorist tradition it grew from.

Strengths

  • Precise and predictive: produces a distinctive, replicable pattern on the cumulative record for every schedule.
  • Wide real-world reach: explains gambling, pay structures, drug-taking and consumer demand, and underpins applied behavior-change programs.

Schedules give operant psychology genuine predictive power. Each one yields a distinctive, replicable signature on the cumulative record, a running graph of total responses over time where a steeper slope means a faster rate.

The signature includes the fixed-ratio break-and-run, the fixed-interval scallop, and the steady variable-ratio rate (Ferster & Skinner, 1957). These patterns recur across species and settings.

The framework predicts the rate and persistence of behavior, not just that reward “works.” That precision matters. It explains behaviorism’s staying power.

The framework also reaches far beyond the laboratory, explaining gambling persistence, pay-structure effects, drug break points and consumer demand. It underpins token economies, programs where desired behaviors earn tokens later exchanged for rewards.

Ayllon and Michael (1959) showed that a psychiatric nurse could act as a “behavioral engineer,” systematically reinforcing adaptive behavior in patients. Their approach launched the token-economy tradition (Morgan, 2010).

A growing body of research since 2015 extends this analysis to smartphones, social media and games. Their variable-reward notification and feed designs mirror slot-machine schedules and may foster compulsive use.

Limitations

  • Cognitive factors are neglected: humans can respond to a described rule rather than their actual reinforcement history.
  • Reductionist and ethically fraught: omits meaning and motivation, and hands institutions a tool for covert behavioral control.
  • Limited animal-to-human generalizability: the core findings come from rats and pigeons in tightly controlled chambers.

The “pure” schedule pattern often breaks down in people. Schedule effects come mainly from controlled animal chambers, but in humans, instructions, expectations and rule-governed behavior can override the contingencies.

A person who understands a schedule may respond to the rule itself rather than to their actual reinforcement history. This limits how directly schedule laws generalize to complex human behavior.

Explaining behavior solely through reinforcement schedules also omits meaning, motivation and cognition, reducing purposive human activity to external contingencies rather than choice.

The same power raises ethical concerns. Employers, institutions and technology companies can arrange schedules that shape behavior without a person’s awareness or consent. Deliberate variable-reward design in gambling and technology is a direct ethical charge against applied schedule engineering.

Schedule effects were also established chiefly with rats and pigeons in tightly controlled chambers. Human behavior is shaped by language, self-awareness and culture, so these “pure” patterns may not transfer directly to people.

Mini Quiz

Below are examples of schedules of reinforcement at work in the real world. Read the examples and then determine which kind of reinforcement schedule is being used.

References

Ayllon, T., & Michael, J. (1959). The psychiatric nurse as a behavioral engineer. Journal of the Experimental Analysis of Behavior, 2(4), 323–334.

Baum, W. M. (1974). On two types of deviation from the matching law: Bias and undermatching. Journal of the Experimental Analysis of Behavior, 22(1), 231–242.

Ferster, C. B., & Skinner, B. F. (1957). Schedules of reinforcement. Appleton-Century-Crofts.

Herrnstein, R. J. (1961). Relative and absolute strength of response as a function of frequency of reinforcement. Journal of the Experimental Analysis of Behavior, 4(3), 267–272.

Herrnstein, R. J. (1970). On the law of effect. Journal of the Experimental Analysis of Behavior, 13(2), 243–266.

Humphreys, L. G. (1939). The effect of random alternation of reinforcement on the acquisition and extinction of conditioned eyelid reactions. Journal of Experimental Psychology, 25(2), 141–158.

Morgan, D. L. (2010). Schedules of reinforcement at 50: A retrospective appreciation. The Psychological Record, 60(1), 151–172.

Myers, David G. (2011). Psychology (10th ed.). Worth Publishers.

Skinner, B. F. (1948). “Superstition” in the pigeon. Journal of Experimental Psychology, 38(2), 168–172. https://doi.org/10.1037/h0055873

Skinner, B. F. (1953). Science and human behavior. Macmillan.

What are schedules of reinforcement?

Schedules of reinforcement are rules that control the timing and frequency of reinforcement delivery in operant conditioning. They include fixed-ratio, variable-ratio, fixed-interval, and variable-interval schedules, each dictating a different pattern of rewards in response to a behavior.

Which schedule of reinforcement is most resistant to the extinction of learned responses?

The variable-ratio schedule of reinforcement is the most resistant to extinction. This is because the reinforcement is given after an unpredictable number of responses, making it more difficult for the behavior to cease. Examples include gambling or lottery games, where a win is unpredictable but can occur anytime.

Saul McLeod, PhD

BSc (Hons) Psychology, MRes, PhD, University of Manchester

Chartered Psychologist (CPsychol)

Saul McLeod, PhD, is a qualified psychology teacher with over 18 years of experience in further and higher education. He has been published in peer-reviewed journals, including the Journal of Clinical Psychology.


Annabelle G.Y. Lim

Psychology Graduate

BA (Hons), Psychology, Harvard University

Annabelle G.Y. Lim is a graduate in psychology from Harvard University. She has served as a research assistant at the Harvard Adolescent Stress & Development Lab.