Learning

Learning is a change in behavior that happens in response to experience, built from a stimulus and the response it produces.

Learning is defined as a change in behavior that happens in response to experience. As we interact with the world and gain new experiences, those experiences change how we behave and how we respond to different situations. Two terms anchor this process: a stimulus is anything in the environment capable of producing a response — something seen, heard, smelled, or felt — and a response is the behavior or reaction that follows it. Touch a hot stove, and the heat is the stimulus while pulling your hand away is the response. The MCAT tests two major categories of learning that build on this idea: associative learning (classical and operant conditioning) and observational learning.

Key Takeaways

  • Learning is a change in behavior resulting from experience, built from a stimulus (something capable of producing a response) and a response (the resulting behavior).

  • Habituation is a gradually decreasing response to a repeated stimulus; dishabituation is that responsiveness returning after something new occurs.

  • Classical conditioning (Pavlov) is stimulus-stimulus learning: a neutral stimulus becomes a conditioned stimulus after repeated pairing with an unconditioned stimulus, eventually producing a conditioned response on its own. Once formed, that association can undergo extinction, spontaneous recovery, generalization, or discrimination.

  • Operant conditioning is behavior-consequence learning: reinforcement increases a behavior, punishment decreases it, and positive/negative describes whether something is added or removed — not whether it's good or bad.

  • The four reinforcement schedules — fixed interval, variable interval, fixed ratio, variable ratio — each depend on a different combination of time/response-count and predictability, and each produces its own distinct response pattern; variable ratio is the most resistant to extinction.

  • Shaping builds complex behaviors by reinforcing successive approximations toward the final goal.

  • Latent learning, preparedness, and instinctive drift show that cognitive and biological factors shape learning beyond simple conditioning — including that learning and performance aren't the same thing.

  • Observational learning (Bandura) happens by watching a model, and depends on four processes: attention, retention, reproduction, and motivation (which can be external, vicarious, or self-generated).

What Is Learning?

One of the simplest ways experience changes behavior is habituation — repeated exposure to the same stimulus causes the response to that stimulus to gradually decrease. Move into an apartment next to a busy road, and the traffic noise might keep you up at first; after living there a while, you stop registering it. The response isn't gone forever, though. If something new or unexpected happens, responsiveness can return — a process called dishabituation. A sudden loud car horn immediately grabs attention again, and afterward the ordinary traffic sounds become noticeable once more.

Associative learning happens when learning occurs by forming associations — either between two different stimuli, or between a behavior and the consequences that follow it. There are two major forms of associative learning tested on the MCAT: classical conditioning and operant conditioning.

Classical Conditioning

Classical conditioning is stimulus-stimulus learning: forming an association between two different stimuli so that one of them gains the ability to produce a response it couldn't produce before. The classic demonstration comes from Ivan Pavlov, who was studying digestion in dogs when he noticed they began salivating not just to food, but to things that regularly happened before the food arrived.

Before conditioning, food naturally and automatically produced salivation — no learning required. Food is the unconditioned stimulus (US), and the salivation it produces is the unconditioned response (UR). Pavlov then introduced the sound of a bell. At this point the bell was a neutral stimulus — it didn't naturally produce salivation.

Pavlov began ringing the bell every time he was about to give the dogs food. At first nothing changed; the dogs still salivated because of the food. But after repeated pairing, the dogs began associating the bell with the food's arrival — a stage of learning called acquisition, the process of forming an association between the neutral stimulus and the unconditioned stimulus. Eventually, ringing the bell alone — without any food — was enough to make the dogs salivate. At that point the bell was no longer neutral; it had become a conditioned stimulus (CS), a previously neutral stimulus that gained the ability to produce a response after repeated pairing with an unconditioned stimulus. The salivation that now occurred in response to the bell alone is the conditioned response (CR) — conditioned because the dogs had to learn it through experience, unlike the naturally occurring unconditioned response.

Once a conditioned association forms, it isn't permanent. Four outcomes are possible:

  • Extinction: the conditioned stimulus is repeatedly presented without the unconditioned stimulus, and the conditioned response gradually decreases. Ring the bell over and over without ever giving food, and the dogs eventually stop salivating to the bell.

  • Spontaneous recovery: even after extinction, the conditioned response can suddenly reappear if some time passes and the conditioned stimulus is presented again — the original association wasn't completely erased.

  • Generalization: stimuli similar to the original conditioned stimulus also produce the conditioned response. A dog conditioned to salivate to one bell might still salivate to a different bell with a similar sound.

  • Discrimination: the ability to distinguish the conditioned stimulus from other similar stimuli — over time, a dog can learn that only one particular bell predicts food, and not salivate to other similar sounds.

Operant Conditioning

Operant conditioning is behavior-consequence learning: it focuses on voluntary behaviors and how the consequences that follow them make those behaviors more or less likely to happen again. Two core concepts drive it: reinforcement, any consequence that increases the likelihood a behavior happens again, and punishment, any consequence that decreases that likelihood. Both reinforcement and punishment can be positive or negative — but here those words don't mean good or bad. Positive means something is being added; negative means something is being removed.

To sort any scenario into one of the four categories, ask two questions: does the behavior become more or less likely (reinforcement or punishment), and is something being added or removed (positive or negative)?

Category

Effect on behavior

What happens

Example

Positive reinforcement

Increases

Something desirable is added

A professor's praise increases effort on future assignments

Negative reinforcement

Increases

Something aversive is removed

Taking medicine removes headache pain, so the behavior repeats

Positive punishment

Decreases

Something aversive is added

A traffic ticket decreases future speeding

Negative punishment

Decreases

Something desirable is removed

Losing recess decreases a child's misbehavior

MCAT Callout — Don't confuse negative reinforcement with punishment: negative reinforcement still increases behavior — the word "negative" only describes what's removed, not whether the outcome is good or bad. Buckling a seatbelt to stop an annoying beeping sound is negative reinforcement (the unpleasant sound is removed, and seatbelt-buckling increases), not punishment.

Reinforcement Schedules

In real life, reinforcement doesn't have to happen every time a behavior occurs. The pattern used to deliver it is called a reinforcement schedule, and schedules differ along two dimensions: whether reinforcement depends on the number of responses (ratio) or the passage of time (interval), and whether it's delivered predictably (fixed) or unpredictably (variable). Combining these gives four schedules, each producing a distinct pattern of behavior.

Schedule

Depends on

Predictability

Response pattern

Fixed interval

Time

Predictable

Scalloped — responding slows right after reinforcement, then speeds up as the next reward approaches

Variable interval

Time

Unpredictable

Moderate, steady rate throughout

Fixed ratio

Number of responses

Predictable

High rate, with a brief pause after each reinforcement

Variable ratio

Number of responses

Unpredictable

Highest and most consistent rate; most resistant to extinction

A rat rewarded once every five minutes for pressing a lever (fixed interval) tends to pause right after getting food, then press more and more frequently as the five minutes elapse — producing the scalloped pattern. A rat rewarded on an unpredictable time delay (variable interval) presses at a steady, moderate rate because it can never predict when the next reward becomes available. A rat rewarded after every ten presses (fixed ratio) presses rapidly, since every press is a step closer to the known reward, but briefly pauses after each one. A rat rewarded after an unpredictable number of presses (variable ratio) presses at the highest and steadiest rate of all four schedules, because any press could be the one that pays off — the same principle behind slot machines, and why variable ratio schedules are the most resistant to extinction.

Shaping

Complex behaviors are rarely performed correctly all at once, so waiting for the complete behavior before reinforcing it might mean it never happens. Shaping is the process of reinforcing successive approximations of a desired behavior — reinforcing small steps that progressively move toward the final goal. Teaching a dog to roll over might start by rewarding it for lying down, then only for turning onto its side, then for rolling farther, until the complete trick has been built from a chain of reinforced smaller steps.

Cognitive and Biological Influences on Learning

Classical and operant conditioning don't explain everything about learning — learning doesn't always require a direct stimulus-stimulus association, and it doesn't always require immediate reinforcement or punishment.

Latent learning is learning that occurs without reinforcement and isn't immediately demonstrated through behavior — the learning stays hidden until there's a reason to demonstrate it. In the classic experiment, rats allowed to explore a maze without any reward still learned its layout; once a reward was later introduced, they navigated the maze far more efficiently, showing that the reward didn't cause the learning — it only gave them a reason to demonstrate learning that had already happened. This shows that learning and performance aren't the same thing: an organism can acquire information without immediately showing a change in observable behavior.

MCAT Callout — Tolman and the cognitive map: this rat-maze experiment comes from psychologist Edward Tolman, whose research in the 1930s introduced latent learning and the related idea of a cognitive map — a mental representation of a maze's (or environment's) layout, built up even without any reward for exploring it.

Cognitive processes matter for problem solving, too — using prior knowledge and reasoning to work toward a solution rather than relying entirely on trial and error. Noticing a horizontal bar on an unfamiliar door and reasoning that it's meant to be pushed, without trying every possible action first, is problem solving in action.

Biology also shapes what's easy or hard to learn. Preparedness is the idea that organisms are biologically predisposed to learn certain associations more easily than others, based on evolutionary history. Humans develop conditioned fears toward things that posed a real evolutionary threat — snakes, spiders — far more easily than toward harmless objects like flowers, even though nothing about direct experience necessarily makes one association easier to form than the other.

MCAT Callout — Seligman and preparedness theory: this concept comes from psychologist Martin Seligman's preparedness theory (1971), which proposed that phobias cluster around ancestrally dangerous stimuli (snakes, heights, spiders) rather than modern hazards (cars, electrical outlets) precisely because of this biological predisposition.

Instinctive drift is a related biological limit on conditioning: an organism gradually reverts to an innate behavior that interferes with a conditioned response. An animal trained through operant conditioning to perform a behavior for food may, over time, have that trained behavior overridden by an instinctive, food-related behavior that conflicts with it.

MCAT Callout — Breland and Breland's "Misbehavior of Organisms": instinctive drift comes from Keller and Marian Breland's 1961 paper of that name — former students of B.F. Skinner who documented animals abandoning trained behaviors (like a raccoon "washing" a coin instead of depositing it) in favor of strong species-specific instincts.

Observational Learning

Observational learning occurs when new information or behaviors are learned by watching other people, rather than through direct personal experience or consequence. The person performing the behavior is the model, and learning by observing and then imitating that behavior is modeling.

One of the most famous demonstrations of observational learning is Albert Bandura's Bobo doll experiment: children who watched an adult model behave aggressively toward an inflatable Bobo doll were significantly more likely to imitate that aggression themselves when later given the opportunity to interact with the doll — strong evidence that behaviors can be learned purely through observation, without any direct conditioning of the child.

Observational learning has also been linked to mirror neurons — neurons that become active both when an individual performs an action and when that individual observes someone else performing the same action. In humans, a broader mirror neuron system has been associated with regions of the frontal and parietal cortex, potentially helping the brain represent others' actions in ways that contribute to imitation.

Watching a behavior doesn't guarantee it will be learned or performed. According to Bandura, observational learning depends on four processes:

  • Attention: paying attention to what the model is doing in the first place. Attention is more likely for models who are admired, interesting, or similar to the observer, or for behavior that's especially unusual or distinctive.

  • Retention: storing the observed behavior in memory so it can be reproduced later, whether hours or much longer after observing it — through a mental image, mental rehearsal, or another form of encoding.

  • Reproduction: being physically capable of actually performing the behavior. Understanding exactly what an expert gymnast did doesn't mean the body is immediately capable of the same movements — that capability can develop with practice.

  • Motivation: having a reason to actually perform the behavior, since learning a behavior and demonstrating it are different things.

Motivation to perform an observed behavior can come from three sources: external reinforcement (expecting a reward), vicarious reinforcement (having observed someone else get rewarded for the behavior), and self-reinforcement (performing the behavior for a sense of personal satisfaction).

Learning new information is only part of the story — once something is learned, it has to be encoded, maintained, and eventually retrieved, which is the subject of Memory.

Why Learning Matters for the MCAT

Learning is one of the most heavily tested content areas in MCAT Psychology/Sociology, usually through passage-based scenarios that describe a behavior change and ask which term or mechanism explains it. Watch for:

  • Negative reinforcement vs. punishment. Negative reinforcement still increases a behavior — it removes something unpleasant. Punishment always decreases a behavior.

  • Ratio vs. interval, fixed vs. variable. Ratio depends on responses, interval on time; fixed is predictable, variable is unpredictable — and each of the four combinations produces its own signature response pattern.

  • Extinction vs. spontaneous recovery. Extinction is the CR fading out; spontaneous recovery is that same CR unexpectedly reappearing after a rest period.

  • Generalization vs. discrimination. Generalization spreads the response to similar stimuli; discrimination narrows it back down to only the original conditioned stimulus.

  • Learning vs. performance. Latent learning demonstrates that an organism can learn something without showing any immediate change in behavior.

Common MCAT Mistakes

  • Calling negative reinforcement a punishment. Negative reinforcement removes something aversive and increases a behavior; punishment always decreases a behavior, regardless of whether something is added or removed.

  • Mixing up ratio/interval with fixed/variable. Ratio and interval describe what triggers reinforcement (responses vs. time); fixed and variable describe how predictable that trigger is. All four combinations are testable, each with its own response pattern.

  • Treating extinction and spontaneous recovery as opposites that can't both apply to the same behavior. They're sequential: a conditioned response can extinguish, then still reappear later as spontaneous recovery — extinction doesn't erase the original association completely.

  • Confusing generalization with discrimination. Generalization is responding to stimuli merely similar to the conditioned stimulus; discrimination is learning to respond only to the original conditioned stimulus and not to similar ones.

MCAT-Style Concept Check

Question: A salesperson closes a sale after an unpredictable number of phone calls — sometimes after 5 calls, sometimes after 20, sometimes after 40. This unpredictability keeps the salesperson dialing at a rapid, sustained rate, rarely pausing for long even during stretches without a sale. Which reinforcement schedule best explains this pattern?

  • A) Fixed interval

  • B) Fixed ratio

  • C) Variable interval

  • D) Variable ratio

Answer: D

Explanation: The reward depends on a number of responses (calls placed), not the passage of time, ruling out (A) and (C), which are both interval schedules. Because the number of calls needed varies unpredictably rather than staying constant, this is a ratio schedule based on responses with unpredictable timing — variable ratio, not fixed ratio (B), which would require a set, predictable number of calls before every sale. Variable ratio schedules produce the highest, steadiest response rate and are the most resistant to extinction, matching the sustained dialing rate described.

FAQ

What's the actual difference between negative reinforcement and punishment?

Negative reinforcement removes something aversive and increases the behavior that removed it — like buckling a seatbelt to stop an annoying beeping sound. Punishment, whether positive (adding something aversive) or negative (removing something desirable), always decreases a behavior. The words "positive" and "negative" describe adding or removing, not good or bad.

Why is variable ratio the hardest reinforcement schedule to extinguish?

Because reinforcement depends on an unpredictable number of responses, there's no way to know which response will pay off — so responding stays high and steady even through long stretches without reinforcement. This unpredictability is also what makes slot machines and similar variable-ratio systems so resistant to extinction.

What does latent learning prove about how learning works?

Latent learning (demonstrated in Tolman's rat-maze experiments) shows that learning can happen without reinforcement and without being immediately visible in behavior — the rats learned the maze's layout during unrewarded exploration and only demonstrated that knowledge once a reward gave them a reason to. It proves learning and performance are separate things.

What are Bandura's four processes required for observational learning?

Attention (noticing what the model does), retention (storing it in memory), reproduction (being physically capable of performing it), and motivation (having a reason to perform it — external, vicarious, or self-generated). All four have to be present for observation to turn into a demonstrated behavior.

More in This Chapter