Prediction Error in Dogs: The Core Mechanism of Learning and Behavior Change
Michael Sauerwein · April 28, 2026
A treat that arrives when none was expected teaches more than the same treat delivered on schedule. A reward that fails to arrive when the dog was certain of it produces something that looks a great deal like disappointment. Both observations point at the same underlying quantity: the gap between what was expected and what occurred.
This article covers prediction error as an established model of learning that is applied to dog training rather than as a canine finding — how it works, how it drives acquisition and extinction, why it produces frustration, and what it implies for reinforcement schedules, cue integrity, and training practice. One framing runs throughout and matters more here than the training literature usually admits. The prediction-error account was established by recording directly from dopamine neurons in behaving monkeys and has since been causally manipulated in rodents. No comparable measurement has been made in dogs. The canine evidence consists of indirect fMRI signals, behavioral experiments, and inference from a well-supported cross-species model. The framework is strong; its canine specifics are extrapolated, and this article marks the line rather than blurring it (the neurochemistry behind it is covered separately).
1. What Prediction Error Is
1.1 The Definition
Prediction error is the difference between the outcome that occurred and the outcome that was expected, where expectation reflects the value estimate the animal has built from prior experience.
A positive prediction error means the outcome was better than predicted — a larger reward than usual, or a reward where none was anticipated. The value of whatever preceded it is revised upward. A negative prediction error means the outcome was worse than predicted — a smaller reward, or none at all. Value is revised downward. A zero prediction error means expectation and outcome matched, and there is nothing to update.
That last case is the counterintuitive one and the practically important one: a perfectly predictable reward delivers no new learning. It maintains behavior; it does not build it.
1.2 It Is Not a Calculation the Dog Performs
Prediction error is not deliberation. It is an automatic signal that determines whether the rest of the system should attend, update, or leave things as they are. A dog does not notice that a reward was better than expected; the discrepancy is registered before anything resembling noticing could occur (which is why behavior does not directly report an internal state).
1.3 The Theoretical Lineage
The idea predates the neuroscience. Rescorla and Wagner (1972) formalized the insight that conditioning depends on how surprising an outcome is rather than on mere pairing — a stimulus already predicted by something else gains little associative strength. Reinforcement learning theory later expressed the same principle as temporal difference learning, in which a value estimate is revised in proportion to the error multiplied by a learning rate (Sutton & Barto, 2018).
1.4 Surprise Alone Does Not Produce Learning
It is tempting to shorten the model to "surprise teaches", and the short version is wrong in a way that matters. A discrepancy produces learning only where the animal can relate it to something: an outcome it can perceive, a predictor it was actually attending to, and a state in which it can process the information at all.
An event that is merely unexpected, with no consequence attached to it, gives the system nothing to update. So does a discrepancy in a dog that is over its working range, or one attached to a feature the handler never intended ([L:/research/arousal-regulation-dogs-neurophysiology-training|where arousal is treated as its own variable]). Surprise is a condition for learning rather than a method of producing it.
1.5 Why the Idea Is Worth the Trouble
Prediction error explains a set of findings that otherwise look unrelated: why timing matters, why a blocked cue is not learned, why intermittent reward is more durable, why extinction is slow. One principle covering four phenomena is why the idea took over learning theory.
That coverage is also the reason to be careful, because a framework that explains a great deal is easy to extend beyond where it was tested (where extinction and relapse are treated in detail). The four phenomena above are within its scope; a good deal of what gets attributed to it is not.
2. Dopamine, Expectation and Prediction Error
2.1 Phasic and Tonic Signaling
Two modes need separating. Phasic dopamine responses are brief, burst-like increases or decreases on a sub-second timescale, and these are what carry the error signal. Tonic dopamine is the sustained background level, which modulates motivation and vigor of pursuit but does not itself encode prediction error. Confusing the two produces the common mistake of treating dopamine as a general "motivation chemical" whose level should simply be raised.
2.2 What the Recordings Showed
Schultz, Dayan and Montague (1997) recorded from midbrain dopamine neurons in monkeys and found the pattern that defines the model: a burst to unpredicted reward, no change when a fully predicted reward arrives on time, and a dip below baseline when a predicted reward is omitted. With learning, the burst migrates backwards to the earliest reliable predictor of reward — which is the mechanistic account of why a marker signal becomes reinforcing.
Waelti, Dickinson and Schultz (2001) supplied the decisive test using a blocking paradigm. Blocking is the classic demonstration that pairing alone is insufficient: a stimulus paired with a reward already fully predicted by another stimulus produces little learning. They showed that dopamine responses tracked prediction error rather than stimulus–reward pairing, and that behavioral learning followed the same pattern. Dopamine responses in that work, in other words, tracked expectation and its violation rather than reward as such.
2.3 Signed and Unsigned Errors
Most of this article concerns signed prediction error, where direction determines whether value goes up or down. Learning models also posit an unsigned error signal encoding the magnitude of surprise irrespective of direction, generally linked to attention and to modulation of the learning rate rather than to value itself. It is less well characterized, and in dogs essentially unstudied.
2.4 A Second Teaching Signal
The picture has recently become more complex. Greenstreet et al. (2025) demonstrated in mice that movement-related dopamine in the tail of the striatum encodes an action prediction error — a value-free signal that reinforces repetition rather than reward, and which consolidates stable stimulus–action associations when paired with reward prediction error circuitry. This is mouse data with causal manipulation, entirely untested in dogs, and it suggests that "dopamine teaches value" is an incomplete summary even in the species where it was established (how habits and automatic behavior are treated more broadly).
2.5 Why the Simple Story Is Under Revision
The account in which dopamine neurons signal a single scalar reward-prediction error has been complicated by recent work showing more heterogeneous signaling (Greenstreet et al.). That revision is happening in the species where the recordings are possible.
A training article that states the simple version confidently is therefore quoting a position the field itself is moving away from. The behavioral predictions are unaffected by that revision, which is the reason to lead with them.
3. What the Canine Evidence Actually Shows
3.1 No Dopamine Recording Exists in Dogs
This deserves stating plainly, because training literature frequently implies otherwise. Nobody has recorded dopamine neuron activity in a dog. There is no canine optogenetics, no microdialysis during learning, no single-unit recording. Every statement about phasic dopamine bursts in a dog's brain is transferred from primates and rodents.
3.2 What Canine Imaging Contributes
Awake canine fMRI provides the closest available evidence, and it is genuinely supportive within its limits. Berns, Brooks and Spivak (2012) trained two dogs to lie still in a scanner and found caudate activation to a hand signal predicting food relative to one predicting nothing. The replication extended this to thirteen dogs and found a positive differential caudate response in eight of thirteen, with a mean differential of 0.09% — comparable to human studies (Berns et al., 2013). Cook et al. (2016) later showed that ventral caudate responses to food- and praise-predicting cues varied between individuals and predicted each dog's subsequent choices.
Three caveats travel with all of it: fMRI measures blood oxygenation rather than dopamine, the samples are small and consist of unusually cooperative scanner-trained dogs, and roughly a third of dogs did not show the group effect at all (the same caveats apply to canine imaging of frontal control).
3.3 The Behavioral Canine Evidence
Behavior offers a more direct route. Reicher et al. (2024) trained 24 family dogs under a controlling and a permissive style and found that dogs experiencing the more rewarding-than-expected condition — a positive expectancy violation — performed better after sleeping. That is a canine demonstration that outcomes exceeding expectation change what is retained, without any claim about the underlying neurochemistry (and sleep does a substantial part of that work).
3.4 What Awake Canine fMRI Can and Cannot Do
It measures blood flow in a scanner-trained dog holding still. That is a genuine achievement and it is several steps removed from recording individual dopamine neurons during learning (Berns, Brooks & Spivak, 2012).
What it can establish is that reward-related regions respond differently to different stimuli and that those differences predict behavior outside the scanner (Cook et al., 2016). What it cannot establish is a prediction-error signal, because that requires resolving single neurons on a timescale fMRI does not reach.
3.5 The Behavioral Evidence Is the Stronger Column
Whether dogs behave as the model predicts is testable without any imaging, and several of the predictions have been examined in this species — blocking, timing effects and schedule effects among them. That evidence carries the practical claims and is rarely the part quoted, because a behavioral result is less impressive to repeat than a brain finding.
An article resting on canine behavior and borrowing primate mechanism is in a better position than one doing the reverse. The reverse construction is common in training material and is what produces confident claims that collapse under checking.
4. What the Model Does Not Explain
4.1 Where the Prediction Comes From
Prediction error is the difference between what was expected and what arrived. What determines the expectation in the first place, and how it is built — how a dog comes to expect anything from a situation it has met a handful of times — the model takes as given.
That is not a flaw; it is the model's scope. It becomes a flaw when the model is presented as a complete account of learning, which is how it usually arrives in training material.
4.2 Which Features the Dog Attends To
The error signal updates the value of whatever the animal treated as the predictor. What the animal treated as the predictor is a question about attention and perception, and the model has nothing to say about it. Two dogs in the same session can update entirely different associations from the same event.
Practically this is where most training failures live: the dog learned something, and it was not the thing the handler thought it was marking. The error signal did its work correctly on the wrong association, which is why the dog appears to have learned nothing when it has in fact learned something else.
4.3 Motivation and State
The same reward produces a different error depending on how hungry, tired, aroused or uncomfortable the animal is. The model handles this by treating reward value as a parameter, which relocates the question rather than answering it (how arousal and state change responses). For a handler the parameter is the whole problem: what this dog will work for, today, in this place.
4.4 Individual Differences
Two dogs in identical conditions learn at different rates, and the model accommodates this with a learning-rate parameter fitted after the fact. That is descriptively useful and explains nothing about why the rates differ (stable differences between individual dogs). A parameter fitted to a result is a summary of the result rather than an account of it.
4.5 Why the Limits Matter Here Specifically
Prediction error is among the most elegant ideas in this collection, and elegance invites overextension. A framework that accounts for timing, blocking, extinction and schedule effects looks as though it should account for everything, and it is quoted accordingly.
Knowing what it leaves out is what keeps it useful, and the omissions above are precisely the variables a practitioner spends most time on: what the dog is attending to, how it is feeling, and why this dog is not the last one.
5. Prediction Error in Operant and Classical Learning
5.1 Operant: Updating the Value of Actions
When a dog performs a behavior and an outcome follows, the value of that behavior is revised by the error. A first-ever reward for sitting produces a large positive error and rapid acquisition. An unexpectedly high-value reward produces a positive error and strengthens the behavior that preceded it. The usual treat, arriving exactly as anticipated, produces approximately nothing — the behavior is maintained, not strengthened.
A conceptual distinction is worth preserving here. The value reduction produced by a negative prediction error is not the same thing as punishment. Punishment is an externally applied consequence; negative prediction error is the system's own updating signal, generated by the absence of an expected outcome. The behavioral result — reduced likelihood — can look similar, but the mechanism and the welfare implications differ.
5.2 Classical: How a Marker Acquires Value
The same logic drives cue learning. Initially a clicker predicts nothing and produces no response. During learning, the click is followed by food and the error is generated at the food. After learning, the response has shifted to the click itself, and the food — now fully predicted — generates no error. The marker has become a conditioned reinforcer precisely because it now occupies the position where the prediction error is produced.
5.3 Extinction
Withholding the reward after an established cue generates a negative prediction error at the moment the reward was due. Repeated, this reduces the cue's expected value. Critically, it does not delete the original association: extinction establishes a new inhibitory learning that competes with the original, which remains available — which is why spontaneous recovery, renewal in a different context, and rapid reacquisition all occur, and why building a competing alternative matters more than erasing the original (the flexibility side of the same process).
5.4 Why Timing Matters So Much
A marker delivered at the moment the behavior occurs identifies which behavior the error signal should update. Delivered two seconds later, it identifies whatever the dog was doing then.
That is the clearest practical consequence of the whole framework, and it follows from the model without requiring any claim about transmitters (what this means for reinforcement timing). A handler who improves nothing except marker timing will usually see a difference, and it costs nothing to try before changing anything else.
6. Negative Prediction Error: Frustration and Poisoned Cues
6.1 The Emotional Cost
A negative prediction error is not an emotionally neutral computation. Unexpected reward omission is associated with negative affect and with activation of regions including the anterior cingulate cortex and amygdala — findings from human and rodent work, extended to dogs by inference rather than measurement.
Behaviorally, the canine picture is clear enough without the neuroscience: dogs with low frustration tolerance respond to unexpected non-reinforcement with escalation, redirection, or disengagement, and for these dogs extinction used alone is a poor tool (the neurobiology of canine frustration in detail).
6.2 Extinction Bursts Are Less Universal Than Commonly Claimed
The transient increase in behavior at the start of extinction is usually presented as inevitable. It is not. Lerman and Iwata (1995) analyzed 113 sets of extinction data and found bursting in 24% of cases — and, importantly for practice, in 36% of cases where extinction was used alone against only 12% where it was combined with other procedures such as reinforcing an alternative behavior.
That is an unusually actionable finding: pairing extinction with differential reinforcement roughly cuts the burst rate. The work is from applied behavior analysis with human participants rather than dogs, so the percentages should not be transferred literally — but the direction of the effect is consistent with the prediction-error account, since an available alternative that still pays keeps positive errors in the picture.
6.3 Poisoned Cues
A cue that has sometimes been followed by something aversive comes to predict a mixture of outcomes. Its expected value drops and the variance of that expectation rises, producing hesitation, conflict, or refusal.
The term is a practitioner concept without a dedicated experimental literature in dogs. The closest hard evidence is Schilder and van der Borg (2004), who found that guard dogs trained with electric shock continued to show stress-related behavior in contexts where no shock occurred, having learned that the handler's presence and commands predicted it. That is the same process operating at the level of the person rather than a single cue (the wider fallout of aversive methods).
6.4 Why Poisoned Cues Are the Best Example in This Article
A cue that sometimes predicts something aversive acquires mixed value, and the dog's response reflects the mixture rather than the intention behind it. That is a prediction the model makes cleanly and that handlers observe constantly.
It also explains why a recall cue used to end enjoyable activity degrades, which is a common and avoidable problem with a mechanism behind it. Calling a dog off a walk only to put it on leash and go home is the standard way to poison the most important cue a household has.
6.5 Negative Errors Are Not a Training Tool
Because negative prediction errors drive learning, it is tempting to manufacture them. The cost is frustration, and the behavioral evidence on frustration in dogs is not encouraging about deliberately producing it.
Managing negative errors rather than maximizing them is the position the evidence supports, and the article says so in its practical chapter. Some negative error is unavoidable in any learning; producing it deliberately is a different decision.
7. Reinforcement Schedules: What Prediction Error Explains and What It Does Not
7.1 The Claim That Needs Correcting
It is widely stated that variable reinforcement produces stronger learning than continuous reinforcement. This conflates two different things. What variable schedules reliably produce is greater resistance to extinction — the partial reinforcement extinction effect — because an animal trained through unrewarded trials finds the absence of reward less surprising when reinforcement stops. That is not the same as faster or better acquisition.
7.2 The Canine Test
Cimarelli et al. (2021) tested this directly. Two groups of naïve dogs were clicker-trained on a novel behavior; one received food after every click, the other after 60% of clicks. Partial rewarding did not improve learning speed — and the partially rewarded dogs subsequently showed a more pessimistic bias in a cognitive bias test than the continuously rewarded group.
So the prediction-error logic that makes variability attractive has a boundary condition, and it is an early one. In a dog that has not yet formed the association, omitted rewards are simply negative prediction errors with an affective cost and no compensating benefit.
7.3 Dogs Do Not Uniformly Prefer the Surprising Option
If unpredictability were reinforcing in itself, dogs offered a choice between a constant and a varied reward should favor the varied one. Bremhorst, Bütler, Würbel and Riemer (2018) ran that test with 16 pet dogs in a concurrent two-choice procedure. At group level the choices did not differ from chance: six dogs preferred the varied reward, six the constant one, and four showed no preference either way.
That is a small sample and a preference test rather than a learning test, and it is still the most direct canine check available on a claim trainers repeat constantly. Variability is a property of the schedule that affects resistance to extinction; it is not something dogs as a species seek out ([L:/research/temperament-personality-coping-styles-dogs|with individual variation doing more work than group means suggest]).
7.4 The Practical Synthesis
Reinforce continuously while a behavior is being acquired. Once it is fluent, introduce variability and occasional larger-than-expected rewards to maintain positive prediction errors and build resistance to extinction. The sequencing is not a stylistic preference; reversing it has been tested in dogs and it did not work (and a dog that cannot access a behavior at all is a different problem again).
7.5 What the Canine Test Contributes
Manipulating partial reward during clicker training in dogs (Cimarelli et al., 2021) tests a prediction the model makes and a recommendation trainers give. That combination — a theoretical claim with a canine test attached — is rare enough in this field to be worth pointing out.
It also produced a result that complicates the standard advice, which is the more informative outcome. A canine test that confirmed the recommendation would have added less.
8. Training with Prediction Error in Mind
8.1 The Handler as a Source of Prediction Error
A handler is part of the environment the dog is predicting. Inconsistent contingencies — a cue that sometimes pays and sometimes does not, criteria that shift without warning, rewards that depend on the handler's mood — generate prediction errors that carry no useful information, because they do not correspond to anything the dog can act on.
There is also direct evidence that human affective state reaches canine decision-making: Parr-Cortes et al. (2024) found that the odor of an unfamiliar stressed person made dogs slower to approach one of three ambiguous locations in a cognitive bias test — eighteen dogs in a single study, but a demonstration that the handler is not a neutral element (how emotional states transfer between species).
8.2 Use Surprise Where It Helps
Occasional jackpots and unexpected reward upgrades produce large positive errors and strengthen what preceded them. Novel reinforcers work for the same reason. This is maintenance-phase technique, not acquisition-phase technique.
8.3 Keep Contingencies Legible
Clear criteria, consistent timing, and a marker used precisely reduce the noise in the dog's predictions. The goal is not the absence of surprise but that surprise carries information.
8.4 Manage Negative Errors Rather Than Maximize Them
Where a behavior needs to reduce, teach an alternative that pays. Differential reinforcement gives the dog a route to positive errors while the old behavior extinguishes, and — per Lerman and Iwata (1995) — substantially reduces the burst. Stepping reward value down gradually before withdrawing it produces smaller errors than an abrupt stop.
The frequently recommended "no reward marker" belongs in a more cautious category. The rationale is coherent: a signal that reinforcement is not coming should reduce uncertainty and speed updating. Whether it does so in dogs without itself acquiring aversive properties has not been established experimentally, and it should be used with that uncertainty in mind.
8.5 Protect Cue Value
Do not follow a cue with an aversive consequence, and avoid giving cues the dog is unlikely to be able to perform. Where a cue has already been compromised, rebuilding requires many positive errors against an established negative history — and retraining an alternative cue is often faster than repairing the damaged one.
9. How to Use This Without Overclaiming
9.1 Say Surprise, Not Dopamine
Every practical point in this article can be made without naming a transmitter. Learning happens when outcomes differ from expectations; a marker works because it arrives before the reward is available; an unpredictable reward is more potent than a predictable one.
None of those require the neurochemistry, and stating them without it avoids a claim the canine evidence does not support. Naming a transmitter adds authority and no information.
9.2 What a Careful Sentence Looks Like
Not "the click triggers a dopamine release" but "the click is informative because it arrives while the outcome is still uncertain". The second is what has been demonstrated behaviorally; the first is an inference from primate recordings, stated as though someone had measured it in the dog standing in front of you.
The difference costs nothing in clarity and a great deal in defensibility, particularly in front of an audience that includes a veterinarian.
9.3 Where the Model Genuinely Helps a Handler
It explains why a cue that always predicts the same thing stops being informative, why an already-predicted reward teaches less, why a second cue added to a reliable one is not learned, and why the moment of marking matters more than the size of the reward.
Those three are counterintuitive, they change what a handler does, and they follow from the behavioral model rather than the neural one. That is as much as any theory needs to earn its place in practice, and more than most manage.
9.4 What to Do When Someone Cites the Neuroscience
Ask which species. The answer is almost always primate or rodent, and the useful follow-up is whether the behavioral prediction has been tested in dogs — which for several of the claims in this article it has. A claim with a canine behavioral test behind it is worth acting on whatever the mechanism turns out to be.
That is a fair question rather than a hostile one, and it separates the claims worth acting on from the ones worth noting. Most people citing the neuroscience are repeating something they read rather than defending a position.
10. Which Findings Come From Which Species
10.1 The Split Is Stark and the Article Says So
The mechanism is primate and rodent neurophysiology, the theory is mathematical psychology, and the canine contribution is imaging and behavior. This article is unusual in stating that in its own third chapter rather than burying it.
10.2 The Mechanism Was Recorded in Monkeys and Confirmed in Rodents
Dopamine neurons signaling the difference between expected and received reward were recorded in primates (Schultz, Dayan & Montague, 1997), and the finding that those responses comply with formal learning theory likewise (Waelti, Dickinson & Schultz, 2001). The causal work establishing that the signal drives learning rather than merely accompanying it was done in rodents, where the necessary manipulations are possible.
Recent work has complicated the simple version considerably (Greenstreet et al.), which matters because the simple version is what circulates in training material.
10.3 The Theory Is Not Empirical at All
The Rescorla-Wagner model is a formal account of associative learning (Rescorla & Wagner, 1972), and the reinforcement-learning framework is computational (Sutton & Barto, 2018). Neither is a finding about any animal; both are frameworks that fit findings.
Citing a textbook of reinforcement learning in a canine article establishes a vocabulary rather than a result, and it is worth noticing which entries in a reference list are of that kind. A long list of references does not distinguish between findings, frameworks and textbooks.
10.4 What Was Measured in Dogs
Awake unrestrained canine fMRI has been established as a method (Berns, Brooks & Spivak, 2012, 2013) and used to predict preferences (Cook et al., 2016). Partial rewarding in clicker training has been manipulated (Cimarelli et al., 2021), preference for constant against varied rewards has been tested directly (Bremhorst et al., 2018), the odor of a stressed human shifts judgment (Parr-Cortes et al., 2024), and shock-collar effects have been documented (Schilder & van der Borg, 2004).
Seven canine sources, none of them recording a dopamine signal, because that has never been done in this species. The invasive recording the primate work required is not something anyone will do to a pet dog, so this gap is permanent rather than pending. That is worth stating plainly, because gaps described as future work imply someone is working on them.
10.5 Why the Practical Chapters Still Stand
The recommendations here — keep contingencies legible, mark precisely, protect cue value, do not manufacture frustration — follow from behavioral findings and would survive a substantial revision of the dopamine account.
That is the test worth applying, and this article passes it more comfortably than most because its practical material was never resting on the neurophysiology. The neuroscience explains why the behavioral findings look as they do and is not what makes them true.
11. Prediction Error Types at a Glance
Positive prediction error — Definition: outcome better than expected. Signal: phasic dopamine burst, established in primates and rodents. Learning effect: strengthens the preceding behavior or cue; drives acquisition. Training use: jackpots, novel reinforcers, variability after fluency.
Zero prediction error — Definition: outcome exactly as predicted. Signal: no net change. Learning effect: no value update; behavior maintained, and repetition may consolidate it through other mechanisms. Training use: maintenance, not building.
Negative prediction error — Definition: outcome worse than expected or absent. Signal: dip below baseline. Learning effect: reduces value; drives extinction. Training use: extinction, ideally combined with differential reinforcement; carries an affective cost.
Unsigned prediction error — Definition: magnitude of surprise regardless of direction. Signal: probably not purely dopaminergic. Learning effect: modulates attention and learning rate rather than value. Training use: conceptual; unstudied in dogs.
Action prediction error — Definition: difference between the action taken and the action predicted. Signal: movement-related dopamine in the tail of the striatum, in mice. Learning effect: value-free reinforcement of repetition; stabilizes habits. Training use: none yet; explanatory only.
12. Research Gaps and Critical Appraisal
The mechanism has never been measured in a dog. Prediction-error coding was established by single-unit recording in monkeys (Schultz et al., 1997; Waelti et al., 2001) and causally confirmed in rodents. There is no canine equivalent, and claims that the model has been "confirmed in dogs" are not supportable.
Canine imaging is indirect and heterogeneous. BOLD signal is not dopamine, samples number in the low tens, and only 8 of 13 dogs showed the differential caudate response in the replication (Berns et al., 2013).
The affective side is largely extrapolated. The association between reward omission and anterior cingulate and amygdala activity comes from human and rodent work (as does most of the canine neurochemistry literature).
Schedule effects are widely misstated. Variable reinforcement increases resistance to extinction; it does not straightforwardly improve acquisition, and in naïve dogs partial rewarding neither sped learning nor left affect unaffected (Cimarelli et al., 2021). Nor do dogs prefer varied rewards as a group: in a two-choice test, six of 16 chose varied, six constant and four neither (Bremhorst et al., 2018).
Extinction bursts are the exception, not the rule. They appeared in 24% of analyzed cases, and in 12% when extinction was combined with other procedures (Lerman & Iwata, 1995) — though that dataset is human applied behavior analysis, not canine.
Poisoned cues and no-reward markers are practitioner constructs. Both are coherent within the prediction-error framework; neither has a dedicated canine experimental literature.
Action prediction error is mouse data. Greenstreet et al. (2025) is strong causal work in one striatal subregion on one task, and its extension to dogs is speculation.
Individual variation is substantial and under-modeled. Reward responsiveness differs markedly between dogs, so the same contingency generates different errors in different animals.
No dopamine recording exists in dogs. The mechanism was recorded in primates (Schultz et al., 1997; Waelti et al., 2001), causally confirmed in rodents, and inferred for this species.
The simple dopamine account is under revision. Recent work reports more heterogeneous signaling than a single scalar error (Greenstreet et al.), in the species where recordings are possible.
Two of the sources are theoretical. The Rescorla-Wagner model and the reinforcement-learning framework are formal accounts rather than findings about animals.
The model does not specify attention. Which feature the dog treats as the predictor determines what gets learned, and the framework takes that as given.
13. Conclusion
Prediction error is the most useful single model available for thinking about how dogs learn, because it explains acquisition, extinction, frustration, cue value, and schedule effects with one mechanism instead of five rules. Its practical content is straightforward: learning happens where outcomes diverge from expectation, so perfectly predictable reinforcement maintains behavior without building it, and unexpected non-reinforcement teaches something at an emotional cost that has to be managed rather than ignored. Its limits should be stated with equal clarity. The dopamine account rests on primate and rodent recordings that have no canine counterpart, the imaging evidence in dogs is indirect and inconsistent across individuals, and at least one popular application of the framework — thinning reinforcement early to strengthen learning — has been tested in dogs and did not survive the test. Used as a way of thinking about expectation rather than as a set of neurochemical claims, the framework earns its place: every marker, every reward, and every withheld reward is information about what the dog should expect next, and training is largely the business of making that information worth having.
Key Insights (Takeaways)
Learning tracks the gap between expectation and outcome, not the pairing itself. Rescorla and Wagner (1972) formalized this, and blocking experiments showed dopamine responses follow prediction error rather than mere stimulus–reward association (Waelti et al., 2001) — which means a fully predicted reward teaches almost nothing.
The dopamine mechanism has never been measured in a dog. It rests on single-unit recording in monkeys and causal work in rodents; canine evidence is fMRI-based, indirect, and inconsistent between individuals, with only 8 of 13 dogs showing the differential caudate response in the replication (Berns et al., 2013).
Variable reinforcement builds resistance to extinction, not faster learning. The two are routinely conflated. In naïve dogs, rewarding 60% of clicks did not improve learning speed and left the dogs with a more pessimistic cognitive bias (Cimarelli et al., 2021), and given the choice, dogs as a group did not prefer a varied reward over a constant one (Bremhorst et al., 2018) — so reinforce continuously during acquisition and introduce variability only once behavior is fluent.
Extinction bursts are not inevitable. They occurred in 24% of analyzed cases, and in only 12% when extinction was combined with reinforcing an alternative behavior (Lerman & Iwata, 1995). Differential reinforcement is the practical lever.
A cue that predicts mixed outcomes loses its value. Shock-trained dogs learned that their handler's presence predicted pain and carried that outside the training context (Schilder & van der Borg, 2004) — the same degradation that "poisoned cue" describes, operating at the level of the person.
References
Berns, G. S., Brooks, A., & Spivak, M. (2013). Replicability and heterogeneity of awake unrestrained canine fMRI responses. PLoS ONE, 8(12), e81698. https://doi.org/10.1371/journal.pone.0081698
Berns, G. S., Brooks, A. M., & Spivak, M. (2012). Functional MRI in awake unrestrained dogs. PLoS ONE, 7(5), e38027. https://doi.org/10.1371/journal.pone.0038027
Bremhorst, A., Bütler, S., Würbel, H., & Riemer, S. (2018). Incentive motivation in pet dogs – preference for constant vs varied food rewards. Scientific Reports, 8, 9756. https://doi.org/10.1038/s41598-018-28079-5
Cimarelli, G., Schoesswender, J., Vitiello, R., Huber, L., & Virányi, Z. (2021). Partial rewarding during clicker training does not improve naïve dogs' learning speed and induces a pessimistic-like affective state. Animal Cognition, 24(1), 107–119. https://doi.org/10.1007/s10071-020-01425-9
Cook, P. F., Prichard, A., Spivak, M., & Berns, G. S. (2016). Awake canine fMRI predicts dogs' preference for praise vs food. Social Cognitive and Affective Neuroscience, 11(12), 1853–1862. https://doi.org/10.1093/scan/nsw102
Greenstreet, F., Martinez Vergara, H., Johansson, Y., Pati, S., Schwarz, L., Lenzi, S. C., Geerts, J. P., Wisdom, M., Gubanova, A., Rollik, L. B., Kaur, J., Moskovitz, T., Cohen, J., Thompson, E., Margrie, T. W., Clopath, C., & Stephenson-Jones, M. (2025). Dopaminergic action prediction errors serve as a value-free teaching signal. Nature, 643(8074), 1333–1342. https://doi.org/10.1038/s41586-025-09008-9
Lerman, D. C., & Iwata, B. A. (1995). Prevalence of the extinction burst and its attenuation during treatment. Journal of Applied Behavior Analysis, 28(1), 93–94. https://doi.org/10.1901/jaba.1995.28-93
Parr-Cortes, Z., Müller, C. T., Talas, L., Mendl, M., Guest, C., & Rooney, N. J. (2024). The odour of an unfamiliar stressed or relaxed person affects dogs' responses to a cognitive bias test. Scientific Reports, 14(1), 15843. https://doi.org/10.1038/s41598-024-66147-1
Reicher, V., Kovács, T., Csibra, B., & Gácsi, M. (2024). Potential interactive effect of positive expectancy violation and sleep on memory consolidation in dogs. Scientific Reports, 14(1), 9487. https://doi.org/10.1038/s41598-024-60166-8
Rescorla, R. A., & Wagner, A. R. (1972). A theory of Pavlovian conditioning: Variations in the effectiveness of reinforcement and nonreinforcement. In A. H. Black & W. F. Prokasy (Eds.), Classical conditioning II: Current research and theory (pp. 64–99). Appleton-Century-Crofts.
Schilder, M. B. H., & van der Borg, J. A. M. (2004). Training dogs with help of the shock collar: Short and long term behavioural effects. Applied Animal Behaviour Science, 85(3–4), 319–334. https://doi.org/10.1016/j.applanim.2003.10.004
Schultz, W., Dayan, P., & Montague, P. R. (1997). A neural substrate of prediction and reward. Science, 275(5306), 1593–1599. https://doi.org/10.1126/science.275.5306.1593
Sutton, R. S., & Barto, A. G. (2018). Reinforcement learning: An introduction (2nd ed.). MIT Press.
Waelti, P., Dickinson, A., & Schultz, W. (2001). Dopamine responses comply with basic assumptions of formal learning theory. Nature, 412(6842), 43–48. https://doi.org/10.1038/35083500