top of page

Michael Sauerwein

Written by

Prediction Error in Dogs: The Core Mechanism of Learning and Behavior Change

A treat that arrives when none was expected teaches more than the same treat delivered on schedule. A reward that fails to arrive when the dog was certain of it produces something that looks a great deal like disappointment. Both observations point at the same underlying quantity: the gap between what was expected and what occurred.


This article covers prediction error as the organizing concept behind reinforcement learning — how it works, how it drives acquisition and extinction, why it produces frustration, and what it implies for reinforcement schedules, cue integrity, and training practice. One framing runs throughout and matters more here than the training literature usually admits. The prediction-error account was established by recording directly from dopamine neurons in behaving monkeys and has since been causally manipulated in rodents. No comparable measurement has been made in dogs. The canine evidence consists of indirect fMRI signals, behavioral experiments, and inference from a well-supported cross-species model. The framework is strong; its canine specifics are extrapolated, and this article marks the line rather than blurring it (the neurochemistry behind it is covered separately).

Golden Retriever sitting on grass raising a paw while focusing on a treat held by a person’s hand in an outdoor setting with a blurred natural background.

1. What Prediction Error Is


1.1 The Definition


Prediction error is the difference between the outcome that occurred and the outcome that was expected, where expectation reflects the value estimate the animal has built from prior experience.


A positive prediction error means the outcome was better than predicted — a larger reward than usual, or a reward where none was anticipated. The value of whatever preceded it is revised upward. A negative prediction error means the outcome was worse than predicted — a smaller reward, or none at all. Value is revised downward. A zero prediction error means expectation and outcome matched, and there is nothing to update.


That last case is the counterintuitive one and the practically important one: a perfectly predictable reward delivers no new learning. It maintains behavior; it does not build it.


1.2 It Is Not a Calculation the Dog Performs


Prediction error is not deliberation. It is an automatic signal that determines whether the rest of the system should attend, update, or leave things as they are. A dog does not notice that a reward was better than expected; the discrepancy is registered before anything resembling noticing could occur (which is why behavior does not directly report an internal state).


1.3 The Theoretical Lineage


The idea predates the neuroscience. Rescorla and Wagner (1972) formalized the insight that conditioning depends on how surprising an outcome is rather than on mere pairing — a stimulus already predicted by something else gains little associative strength. Reinforcement learning theory later expressed the same principle as temporal difference learning, in which a value estimate is revised in proportion to the error multiplied by a learning rate (Sutton & Barto, 2018).



2. Dopamine as the Teaching Signal


2.1 Phasic and Tonic Signalling


Two modes need separating. Phasic dopamine responses are brief, burst-like increases or decreases on a sub-second timescale, and these are what carry the error signal. Tonic dopamine is the sustained background level, which modulates motivation and vigour of pursuit but does not itself encode prediction error. Confusing the two produces the common mistake of treating dopamine as a general "motivation chemical" whose level should simply be raised.


2.2 What the Recordings Showed


Schultz, Dayan and Montague (1997) recorded from midbrain dopamine neurons in monkeys and found the pattern that defines the model: a burst to unpredicted reward, no change when a fully predicted reward arrives on time, and a dip below baseline when a predicted reward is omitted. With learning, the burst migrates backwards to the earliest reliable predictor of reward — which is the mechanistic account of why a marker signal becomes reinforcing.


Waelti, Dickinson and Schultz (2001) supplied the decisive test using a blocking paradigm. Blocking is the classic demonstration that pairing alone is insufficient: a stimulus paired with a reward already fully predicted by another stimulus produces little learning. They showed that dopamine responses tracked prediction error rather than stimulus–reward pairing, and that behavioral learning followed the same pattern. Dopamine, in other words, behaves the way formal learning theory says a teaching signal should.


2.3 Signed and Unsigned Errors


Most of this article concerns signed prediction error, where direction determines whether value goes up or down. Learning models also posit an unsigned error signal encoding the magnitude of surprise irrespective of direction, generally linked to attention and to modulation of the learning rate rather than to value itself. It is less well characterized, and in dogs essentially unstudied.


2.4 A Second Teaching Signal


The picture has recently become more complex. Greenstreet et al. (2025) demonstrated in mice that movement-related dopamine in the tail of the striatum encodes an action prediction error — a value-free signal that reinforces repetition rather than reward, and which consolidates stable stimulus–action associations when paired with reward prediction error circuitry. This is mouse data with causal manipulation, entirely untested in dogs, and it suggests that "dopamine teaches value" is an incomplete summary even in the species where it was established (how habits and automatic behavior are treated more broadly).



3. What the Canine Evidence Actually Shows


3.1 No Dopamine Recording Exists in Dogs


This deserves stating plainly, because training literature frequently implies otherwise. Nobody has recorded dopamine neuron activity in a dog. There is no canine optogenetics, no microdialysis during learning, no single-unit recording. Every statement about phasic dopamine bursts in a dog's brain is transferred from primates and rodents.


3.2 What Canine Imaging Contributes


Awake canine fMRI provides the closest available evidence, and it is genuinely supportive within its limits. Berns, Brooks and Spivak (2012) trained two dogs to lie still in a scanner and found caudate activation to a hand signal predicting food relative to one predicting nothing. The replication extended this to thirteen dogs and found a positive differential caudate response in eight of thirteen, with a mean differential of 0.09% — comparable to human studies (Berns et al., 2013). Cook et al. (2016) later showed that ventral caudate responses to food- and praise-predicting cues varied between individuals and predicted each dog's subsequent choices.


Three caveats travel with all of it: fMRI measures blood oxygenation rather than dopamine, the samples are small and consist of unusually cooperative scanner-trained dogs, and roughly a third of dogs did not show the group effect at all (the same caveats apply to canine imaging of frontal control).


3.3 The Behavioral Canine Evidence


Behavior offers a more direct route. Reicher et al. (2024) trained 24 family dogs under a controlling and a permissive style and found that dogs experiencing the more rewarding-than-expected condition — a positive expectancy violation — performed better after sleeping. That is a canine demonstration that outcomes exceeding expectation change what is retained, without any claim about the underlying neurochemistry (and sleep does a substantial part of that work).



4. Prediction Error in Operant and Classical Learning


4.1 Operant: Updating the Value of Actions


When a dog performs a behavior and an outcome follows, the value of that behavior is revised by the error. A first-ever reward for sitting produces a large positive error and rapid acquisition. An unexpectedly high-value reward produces a positive error and strengthens the behavior that preceded it. The usual treat, arriving exactly as anticipated, produces approximately nothing — the behavior is maintained, not strengthened.


A conceptual distinction is worth preserving here. The value reduction produced by a negative prediction error is not the same thing as punishment. Punishment is an externally applied consequence; negative prediction error is the system's own updating signal, generated by the absence of an expected outcome. The behavioral result — reduced likelihood — can look similar, but the mechanism and the welfare implications differ.


4.2 Classical: How a Marker Acquires Value


The same logic drives cue learning. Initially a clicker predicts nothing and produces no response. During learning, the click is followed by food and the error is generated at the food. After learning, the response has shifted to the click itself, and the food — now fully predicted — generates no error. The marker has become a conditioned reinforcer precisely because it now occupies the position where the prediction error is produced.


4.3 Extinction


Withholding the reward after an established cue generates a negative prediction error at the moment the reward was due. Repeated, this reduces the cue's expected value. Critically, it does not delete the original association: extinction establishes a new inhibitory learning that competes with the original, which remains available — which is why spontaneous recovery, renewal in a different context, and rapid reacquisition all occur (the full account of why extinguished behavior returns), and why building a competing alternative matters more than erasing the original (the flexibility side of the same process).



5. Negative Prediction Error: Frustration and Poisoned Cues


5.1 The Emotional Cost


A negative prediction error is not an emotionally neutral computation. Unexpected reward omission is associated with negative affect and with activation of regions including the anterior cingulate cortex and amygdala — findings from human and rodent work, extended to dogs by inference rather than measurement.


Behaviorally, the canine picture is clear enough without the neuroscience: dogs with low frustration tolerance respond to unexpected non-reinforcement with escalation, redirection, or disengagement, and for these dogs extinction used alone is a poor tool (the neurobiology of canine frustration in detail).


5.2 Extinction Bursts Are Less Universal Than Commonly Claimed


The transient increase in behavior at the start of extinction is usually presented as inevitable. It is not. Lerman and Iwata (1995) analysed 113 sets of extinction data and found bursting in 24% of cases — and, importantly for practice, in 36% of cases where extinction was used alone against only 12% where it was combined with other procedures such as reinforcing an alternative behavior.


That is an unusually actionable finding: pairing extinction with differential reinforcement roughly cuts the burst rate. The work is from applied behavior analysis with human participants rather than dogs, so the percentages should not be transferred literally — but the direction of the effect is consistent with the prediction-error account, since an available alternative that still pays keeps positive errors in the picture.


5.3 Poisoned Cues


A cue that has sometimes been followed by something aversive comes to predict a mixture of outcomes. Its expected value drops and the variance of that expectation rises, producing hesitation, conflict, or refusal.


The term is a practitioner concept without a dedicated experimental literature in dogs. The closest hard evidence is Schilder and van der Borg (2004), who found that guard dogs trained with electric shock continued to show stress-related behavior in contexts where no shock occurred, having learned that the handler's presence and commands predicted it. That is the same process operating at the level of the person rather than a single cue (the wider fallout of aversive methods).



6. Reinforcement Schedules: What Prediction Error Explains and What It Does Not


6.1 The Claim That Needs Correcting


It is widely stated that variable reinforcement produces stronger learning than continuous reinforcement. This conflates two different things. What variable schedules reliably produce is greater resistance to extinction — the partial reinforcement extinction effect — because an animal trained through unrewarded trials finds the absence of reward less surprising when reinforcement stops. That is not the same as faster or better acquisition.


6.2 The Canine Test

Cimarelli et al. (2021) tested this directly. Two groups of naïve dogs were clicker-trained on a novel behavior; one received food after every click, the other after 60% of clicks. Partial rewarding did not improve learning speed — and the partially rewarded dogs subsequently showed a more pessimistic bias in a cognitive bias test than the continuously rewarded group.


So the prediction-error logic that makes variability attractive has a boundary condition, and it is an early one. In a dog that has not yet formed the association, omitted rewards are simply negative prediction errors with an affective cost and no compensating benefit (schedules and their effects in full).


6.3 The Practical Synthesis


Reinforce continuously while a behavior is being acquired. Once it is fluent, introduce variability and occasional larger-than-expected rewards to maintain positive prediction errors and build resistance to extinction. The sequencing is not a stylistic preference; reversing it has been tested in dogs and it did not work (and a dog that cannot access a behavior at all is a different problem again).



7. Training with Prediction Error in Mind


7.1 The Handler as a Source of Prediction Error


A handler is part of the environment the dog is predicting. Inconsistent contingencies — a cue that sometimes pays and sometimes does not, criteria that shift without warning, rewards that depend on the handler's mood — generate prediction errors that carry no useful information, because they do not correspond to anything the dog can act on.


There is also direct evidence that human affective state reaches canine decision-making: Parr-Cortes et al. (2024) found that the odour of an unfamiliar stressed person made dogs slower to approach one of three ambiguous locations in a cognitive bias test — eighteen dogs in a single study, but a demonstration that the handler is not a neutral element (how emotional states transfer between species).


7.2 Use Surprise Where It Helps


Occasional jackpots and unexpected reward upgrades produce large positive errors and strengthen what preceded them. Novel reinforcers work for the same reason. This is maintenance-phase technique, not acquisition-phase technique.


7.3 Keep Contingencies Legible


Clear criteria, consistent timing, and a marker used precisely reduce the noise in the dog's predictions. The goal is not the absence of surprise but that surprise carries information (which is also what keeps arousal in a workable range).


7.4 Manage Negative Errors Rather Than Maximize Them


Where a behavior needs to reduce, teach an alternative that pays. Differential reinforcement gives the dog a route to positive errors while the old behavior extinguishes, and — per Lerman and Iwata (1995) — substantially reduces the burst. Stepping reward value down gradually before withdrawing it produces smaller errors than an abrupt stop.


The frequently recommended "no reward marker" belongs in a more cautious category. The rationale is coherent: a signal that reinforcement is not coming should reduce uncertainty and speed updating. Whether it does so in dogs without itself acquiring aversive properties has not been established experimentally, and it should be used with that uncertainty in mind (and observed against measurable criteria rather than impression).


7.5 Protect Cue Value


Do not follow a cue with an aversive consequence, and avoid giving cues the dog is unlikely to be able to perform. Where a cue has already been compromised, rebuilding requires many positive errors against an established negative history — and retraining an alternative cue is often faster than repairing the damaged one.



8. Summary: Prediction Error Types at a Glance


Positive prediction error — Definition: outcome better than expected. Signal: phasic dopamine burst, established in primates and rodents. Learning effect: strengthens the preceding behavior or cue; drives acquisition. Training use: jackpots, novel reinforcers, variability after fluency.


Zero prediction error — Definition: outcome exactly as predicted. Signal: no net change. Learning effect: no value update; behavior maintained, and repetition may consolidate it through other mechanisms. Training use: maintenance, not building.


Negative prediction error — Definition: outcome worse than expected or absent. Signal: dip below baseline. Learning effect: reduces value; drives extinction. Training use: extinction, ideally combined with differential reinforcement; carries an affective cost.


Unsigned prediction error — Definition: magnitude of surprise regardless of direction. Signal: probably not purely dopaminergic. Learning effect: modulates attention and learning rate rather than value. Training use: conceptual; unstudied in dogs.


Action prediction error — Definition: difference between the action taken and the action predicted. Signal: movement-related dopamine in the tail of the striatum, in mice. Learning effect: value-free reinforcement of repetition; stabilizes habits. Training use: none yet; explanatory only.



9. Research Gaps and Critical Appraisal


The mechanism has never been measured in a dog. Prediction-error coding was established by single-unit recording in monkeys (Schultz et al., 1997; Waelti et al., 2001) and causally confirmed in rodents. There is no canine equivalent, and claims that the model has been "confirmed in dogs" are not supportable.


Canine imaging is indirect and heterogeneous. BOLD signal is not dopamine, samples number in the low tens, and only 8 of 13 dogs showed the differential caudate response in the replication (Berns et al., 2013).


The affective side is largely extrapolated. The association between reward omission and anterior cingulate and amygdala activity comes from human and rodent work (as does most of the canine neurochemistry literature).


Schedule effects are widely misstated. Variable reinforcement increases resistance to extinction; it does not straightforwardly improve acquisition, and in naïve dogs partial rewarding neither sped learning nor left affect unaffected (Cimarelli et al., 2021).


Extinction bursts are the exception, not the rule. They appeared in 24% of analysed cases, and in 12% when extinction was combined with other procedures (Lerman & Iwata, 1995) — though that dataset is human applied behavior analysis, not canine.


Poisoned cues and no-reward markers are practitioner constructs. Both are coherent within the prediction-error framework; neither has a dedicated canine experimental literature.


Action prediction error is mouse data. Greenstreet et al. (2025) is strong causal work in one striatal subregion on one task, and its extension to dogs is speculation.


Individual variation is substantial and under-modelled. Reward responsiveness differs markedly between dogs, so the same contingency generates different errors in different animals (as temperament research would predict).



10. Conclusion


Prediction error is the most useful single idea available for thinking about how dogs learn, because it explains acquisition, extinction, frustration, cue value, and schedule effects with one mechanism instead of five rules. Its practical content is straightforward: learning happens where outcomes diverge from expectation, so perfectly predictable reinforcement maintains behavior without building it, and unexpected non-reinforcement teaches something at an emotional cost that has to be managed rather than ignored. Its limits should be stated with equal clarity. The dopamine account rests on primate and rodent recordings that have no canine counterpart, the imaging evidence in dogs is indirect and inconsistent across individuals, and at least one popular application of the framework — thinning reinforcement early to strengthen learning — has been tested in dogs and did not survive the test. Used as a way of thinking about expectation rather than as a set of neurochemical claims, the framework earns its place: every marker, every reward, and every withheld reward is information about what the dog should expect next, and training is largely the business of making that information worth having.



Key Insights (Takeaways)


  • Learning tracks the gap between expectation and outcome, not the pairing itself. Rescorla and Wagner (1972) formalized this, and blocking experiments showed dopamine responses follow prediction error rather than mere stimulus–reward association (Waelti et al., 2001) — which means a fully predicted reward teaches almost nothing.

  • The dopamine mechanism has never been measured in a dog. It rests on single-unit recording in monkeys and causal work in rodents; canine evidence is fMRI-based, indirect, and inconsistent between individuals, with only 8 of 13 dogs showing the differential caudate response in the replication (Berns et al., 2013).

  • Variable reinforcement builds resistance to extinction, not faster learning. The two are routinely conflated. In naïve dogs, rewarding 60% of clicks did not improve learning speed and left the dogs with a more pessimistic cognitive bias (Cimarelli et al., 2021) — so reinforce continuously during acquisition and introduce variability only once behavior is fluent.

  • Extinction bursts are not inevitable. They occurred in 24% of analysed cases, and in only 12% when extinction was combined with reinforcing an alternative behavior (Lerman & Iwata, 1995). Differential reinforcement is the practical lever.

  • A cue that predicts mixed outcomes loses its value. Shock-trained dogs learned that their handler's presence predicted pain and carried that outside the training context (Schilder & van der Borg, 2004) — the same degradation that "poisoned cue" describes, operating at the level of the person.



References


Berns, G. S., Brooks, A., & Spivak, M. (2013). Replicability and heterogeneity of awake unrestrained canine fMRI responses. PLoS ONE, 8(12), e81698. https://doi.org/10.1371/journal.pone.0081698


Berns, G. S., Brooks, A. M., & Spivak, M. (2012). Functional MRI in awake unrestrained dogs. PLoS ONE, 7(5), e38027. https://doi.org/10.1371/journal.pone.0038027


Cimarelli, G., Schoesswender, J., Vitiello, R., Huber, L., & Virányi, Z. (2021). Partial rewarding during clicker training does not improve naïve dogs' learning speed and induces a pessimistic-like affective state. Animal Cognition, 24(1), 107–119. https://doi.org/10.1007/s10071-020-01425-9


Cook, P. F., Prichard, A., Spivak, M., & Berns, G. S. (2016). Awake canine fMRI predicts dogs' preference for praise vs food. Social Cognitive and Affective Neuroscience, 11(12), 1853–1862. https://doi.org/10.1093/scan/nsw102


Greenstreet, F., Martinez Vergara, H., Johansson, Y., Pati, S., Schwarz, L., Lenzi, S. C., Geerts, J. P., Wisdom, M., Gubanova, A., Rollik, L. B., Kaur, J., Moskovitz, T., Cohen, J., Thompson, E., Margrie, T. W., Clopath, C., & Stephenson-Jones, M. (2025). Dopaminergic action prediction errors serve as a value-free teaching signal. Nature, 643(8074), 1333–1342. https://doi.org/10.1038/s41586-025-09008-9


Lerman, D. C., & Iwata, B. A. (1995). Prevalence of the extinction burst and its attenuation during treatment. Journal of Applied Behavior Analysis, 28(1), 93–94. https://doi.org/10.1901/jaba.1995.28-93


Parr-Cortes, Z., Müller, C. T., Talas, L., Mendl, M., Guest, C., & Rooney, N. J. (2024). The odour of an unfamiliar stressed or relaxed person affects dogs' responses to a cognitive bias test. Scientific Reports, 14(1), 15843. https://doi.org/10.1038/s41598-024-66147-1


Reicher, V., Kovács, T., Csibra, B., & Gácsi, M. (2024). Potential interactive effect of positive expectancy violation and sleep on memory consolidation in dogs. Scientific Reports, 14(1), 9487. https://doi.org/10.1038/s41598-024-60166-8


Rescorla, R. A., & Wagner, A. R. (1972). A theory of Pavlovian conditioning: Variations in the effectiveness of reinforcement and nonreinforcement. In A. H. Black & W. F. Prokasy (Eds.), Classical conditioning II: Current research and theory (pp. 64–99). Appleton-Century-Crofts.


Schilder, M. B. H., & van der Borg, J. A. M. (2004). Training dogs with help of the shock collar: Short and long term behavioural effects. Applied Animal Behaviour Science, 85(3–4), 319–334. https://doi.org/10.1016/j.applanim.2003.10.004


Schultz, W., Dayan, P., & Montague, P. R. (1997). A neural substrate of prediction and reward. Science, 275(5306), 1593–1599. https://doi.org/10.1126/science.275.5306.1593


Sutton, R. S., & Barto, A. G. (2018). Reinforcement learning: An introduction (2nd ed.). MIT Press.

Waelti, P., Dickinson, A., & Schultz, W. (2001). Dopamine responses comply with basic assumptions of formal learning theory. Nature, 412(6842), 43–48. https://doi.org/10.1038/35083500

24. April 2026

bottom of page
unterHUNDs.de Trainerausbildung Blog