Cognitive Bias in Dogs: Measuring Optimism, Pessimism and Mood
Michael Sauerwein · July 25, 2026
A bowl sits halfway between two places a dog knows well. On the left, bowls have always contained food. On the right, they have always been empty. The dog has never seen a bowl in the middle before. How fast it walks over is the entire measurement — and for twenty years it has been the best available answer to a question that otherwise has none.
That question is whether an animal is in a good or bad mood, and the reason it matters is that behavior alone will not tell you. This article covers the logic behind the test, how it works in practice, what the canine studies have actually found — including that the body odor of a stressed stranger is enough to change how a dog responds — and what happened when researchers tried to shift a dog's bias deliberately. The method has produced meaningful results, and it has also produced null results and unresolved methodological problems, both of which get their own sections here. The honest summary is that the effect is real, small, and considerably noisier than the popular version suggests (the general problem of inferring emotion from behavior).
1. The Measurement Problem
1.1 Why Behavior Is Not Enough
A dog that has stopped struggling may be calm or may have given up. A dog with a low tail may be relaxed or worried. The same visible behavior can arise from different emotional states, and the same emotional state can produce different behaviors depending on context. Physiological measures do not resolve this either: cortisol moves with arousal rather than valence, so it rises under excitement as readily as under distress (which is why operational definitions matter).
The gap this leaves is specific: there is still no validated indicator of positive emotional state in dogs (Csoltova & Mehinagic, 2020), and the standard indicator set differentiates arousal only under negative valence (Flint et al., 2024). Absence of visible distress is not evidence of wellbeing.
1.2 The Idea Borrowed From Human Psychology
In humans, mood biases judgment. Anxious and depressed people interpret ambiguous information more negatively — the same neutral face, the same ambiguous sentence, read as threatening rather than benign. This is well established and, crucially, it is measurable through choices rather than through self-report.
If the mechanism is general rather than specifically human, then an animal in a negative mood should also treat ambiguity more cautiously. That is a testable claim, and testing it does not require the animal to say anything.
1.3 The Founding Experiment
Harding, Paul and Mendl (2004) trained rats to press a lever for food when they heard one tone and to avoid pressing when they heard another that signaled an unpleasant noise. Once the discrimination was solid, they played tones of intermediate pitch that the animals had never encountered.
Rats housed in unpredictable conditions responded to the ambiguous tones as though they predicted the bad outcome. The finding launched an entire field, and the underlying framework was developed further into a general account of animal emotion and mood (Mendl, Burman, Parker & Paul, 2009; Mendl, Burman & Paul, 2010b).
1.4 Why This Article Sits Under the Others
Several topics in this collection report that something changed a dog's emotional state, and a good proportion of those findings were produced with this method. Understanding what it measures therefore affects how much weight those findings can carry.
That is an unusual position for a single paradigm to occupy and a reason to examine it carefully rather than to take the results at face value.
1.5 What the Method Is Not
It is not a mood ring, not a diagnostic test, and not something a trainer can run in a session. It is a research instrument that produces a group-level comparison under controlled conditions.
Every practical claim in this article follows from that description, and the claims that circulate about it generally do not.
2. How Cognitive Bias Tests Measure Dog Emotion
2.1 The Spatial Version Used With Dogs
Nearly all canine work uses a spatial task. A bowl in one location always contains food; a bowl in another location is always empty. Once the dog reliably runs to the first and hesitates at the second, bowls appear in ambiguous locations between the two — typically three: near-positive, middle, and near-negative.
The measurement is latency: how quickly the dog approaches the ambiguous bowl. Faster approach is read as an "optimistic" response — the dog behaving as though food is likely. Slower approach is read as "pessimistic."
2.2 What the Terms Actually Mean
"Optimistic" and "pessimistic" here are technical labels for response patterns, not claims about canine belief or personality. A "pessimistic" response does not mean the dog expects the worst; it means the ambiguity was resolved in a direction statistically associated with a negative affective state. The dog is not making a prediction about its life — it is responding to a cue whose meaning it cannot settle, and the direction of that response is used as a proxy.
The distinction matters because the everyday sense of the words is what most secondary coverage runs with, and it overstates what the data supports.
2.3 What Else Could Produce the Same Latency
Affective state is the hypothesis the paradigm was built to test, and it is not the only thing that could move an approach time. Attention, motivation on the day, learning history, and what the dog has come to expect from this room and this experimenter all feed into the same number.
The literature treats those as sources of noise to be controlled, and they are also candidate explanations in their own right. Krahn et al. (2024) is the clearest demonstration: what shifted the score there was prior discrimination training, not mood ([L:/research/operationalizing-dog-behavior-scientific-measurement|which is what operationalizing a behavioral claim is for]). A result is evidence about affective state to the extent that these alternatives were held constant, and no further.
2.4 Why Ambiguity Is the Whole Point
Trained cues tell you only that the dog learned the discrimination. The untrained middle cue is where prior expectation has to fill the gap, and prior expectation is what mood is hypothesized to shift. The test is therefore a controlled way of asking what the dog expects when the evidence runs out (which connects it to prediction error more broadly).
2.5 Why Latency Rather Than Choice
The dog is not asked to choose between locations; it is timed on how quickly it approaches an ambiguous one. That design choice matters, because latency is a continuous measure and produces a gradient rather than a binary.
It also imports every other influence on how fast a dog moves, which is the cost of the choice and the subject of a later chapter (arousal as one such influence).
2.6 The Terms Are Shorthand
Optimistic and pessimistic are labels for faster and slower approach to ambiguity, and the accurate long form is that the dog evaluated an ambiguous situation more or less negatively on that occasion. They are not claims about a dog's outlook, and the literature uses them because they are memorable rather than because they are accurate.
A reader encountering a headline about optimistic dogs is encountering a latency difference with a word attached to it.
3. What the Canine Studies Show
3.1 Separation-Related Behavior
One of the earliest and most influential canine studies tested dogs in a rehoming kennel and compared those that showed separation-related behavior when left alone with those that did not. Dogs displaying separation-related behavior responded to ambiguous bowl locations more slowly — a pessimistic pattern consistent with a negative underlying affective state (Mendl et al., 2010a).
The interpretive value is real. It provides evidence that separation-related behavior is not a training failure or a bid for attention but sits on top of a genuinely negative state (as the neurobiology of that presentation indicates).
3.2 Training Methods
Dogs whose owners reported using two or more aversive training methods approached ambiguous locations more slowly than dogs trained with reward-based methods (Casey et al., 2021). The design was cross-sectional and the training data owner-reported, so causation is not established — but the result converges with independent welfare findings (on the neurological effects of aversive methods).
3.3 Chronic Conditions and Pain
Cavalier King Charles Spaniels with syringomyelia showed response patterns consistent with a negative affective state (Cockburn et al., 2018), and dogs with epilepsy differed from controls on judgment and attention bias measures (Hobbs et al., 2020). Both fit a wider pattern in which physical conditions shift emotional state in ways that are easy to miss clinically (as with visceral pain).
3.4 Housing and Age
Shelter dogs behaved more pessimistically than pet dogs in one carefully analyzed comparison of 51 shelter and 40 pet dogs — though, importantly, only under one method of scoring (Burani et al., 2020). Age also affects performance on discrimination, reversal and cognitive bias tasks in family dogs (Piotti et al., 2018), which means age has to be controlled rather than assumed away (and cognitive decline complicates it further).
3.5 What the Canine Findings Have in Common
Across separation-related behavior, training method, pain, housing and age, the pattern is consistent: conditions expected to produce a worse affective state are associated with slower approach to ambiguous cues.
That consistency across unrelated questions is a stronger argument for the method's validity than any single result, because the studies were run by different groups asking different questions.
3.6 Pain Deserves Separate Mention
Negative judgment bias in dogs with chronic pain and with neurological conditions (Cockburn et al., 2018; Hobbs et al., 2020) is among the more consequential findings here, because it supplies a measurable link between a physical state and an affective one.
It also supports the recurring recommendation across this collection: investigate the body before attributing a behavioral change to something else.
4. What Changes Cognitive Bias in Dogs
4.1 Treatment Changes It
The most clinically important canine result concerns a question that behavior outcome measures cannot answer. When a dog stops howling and stops destroying the door frame, two very different things may have happened: the dog may feel better, or the dog may simply have stopped doing the thing the owner complained about. Owner-reported improvement cannot distinguish them, and neither can a checklist of observed behaviors.
Dogs with separation-related problems treated with fluoxetine alongside a behavior modification plan showed a shift toward less pessimistic responding over the course of treatment (Karagiannis, Burman & Mills, 2015). The authors framed it as the first demonstration that clinical treatment of a negative affective state in a non-human species produces a measurable shift in cognitive bias.
That is a substantial claim, and its significance is methodological rather than pharmacological. It establishes that the question did the animal actually get better? is answerable in principle, using a measure that does not depend on the owner's judgment or on the behavior that prompted the referral (which is exactly the distinction suppression obscures). One study without an untreated control group cannot settle how well the method performs — but it sets the standard any outcome claim should be measured against.
4.2 Another Person's Stress Changes It
Dogs exposed to body odor collected from an unfamiliar stressed person responded more pessimistically to ambiguous bowl locations than dogs exposed to odor from a relaxed person (Parr-Cortes et al., 2024). The person was a stranger the dogs never saw.
Two implications follow. Chemical signals of human stress are detectable and affect canine decision-making (through olfactory channels), and the transfer of emotional state between humans and dogs is not limited to visible or familiar cues (the emotional contagion picture).
4.3 What This Adds Up To
The manipulations that move canine judgment bias are the ones a welfare framework would predict: chronic stress, aversive handling, unresolved separation distress, physical illness, and proximity to a stressed human. Nothing in the canine literature contradicts the framework. The question is how reliably any single test detects it.
What the test supplies is an indication about evaluation patterns rather than a verdict on welfare. A group that evaluates ambiguity more negatively is a group worth looking at more closely, which is a different claim from having established that its welfare is poor ([L:/research/behavior-does-not-equal-emotion-in-dogss|since the observable and the inferred are routinely conflated]).
4.4 Why the Treatment Study Carries Weight
Measuring a shift in judgment bias during treatment (Karagiannis, Burman & Mills, 2015) is more informative than a cross-sectional comparison, because it uses each dog as its own control and because the endpoint does not depend on the owner's impression.
Its limitation is the sample: five dogs in the treatment group, and a combined intervention that cannot separate the medication from the behavior plan.
4.5 The Odor Finding Is the Most Surprising
That the odor of an unfamiliar stressed human shifts a dog's judgment bias (Parr-Cortes et al., 2024) is the result in this article least likely to have been predicted, and the one with the widest practical implications.
It means a variable nobody controls — the state of a person the dog has never met — influences a measure of the dog's affective state. For consultations, veterinary visits and training halls, that is worth knowing.
5. The Null Results
5.1 Brief Owner Absence Did Nothing
Pet dogs briefly separated from their owners showed no negative judgment bias afterwards (Müller et al., 2012). The manipulation was mild and short, and the result is best read as a boundary condition: the test does not respond to every stressor, and a brief separation in a familiar setting is not equivalent to the chronic separation distress in the kennel study.
5.2 Why Null Results Deserve Space Here
A field where positive findings are widely reported and null findings are not produces a distorted picture, and this literature has that shape. Null results define where a method works. Reporting only the successes leaves practitioners believing the test is more sensitive than it is (the same problem that affects reconsolidation research).
5.3 What the Null Result Actually Rules Out
Brief owner absence not shifting judgment bias (Müller et al., 2012) does not show that dogs are unaffected by being left. It shows that this measure did not detect a difference under those conditions, with that sample and that duration.
Those are different claims, and the second is the one the study supports. A measure failing to detect something is evidence about the measure as much as about the dogs.
5.4 Why Including It Matters
An article that reported only the positive findings would give a misleading impression of how reliably the method detects things. The null results are part of the picture of what the test can do.
They also constrain the practical claims: a measure that missed an effect of brief separation is not a measure that will detect small welfare differences in a household.
6. Limits of Cognitive Bias Tests in Dogs
6.1 The Cross-Species Effect Is Small
The largest systematic review and meta-analysis of judgment bias across species found that non-pharmacological manipulations of affect do alter judgment bias in the expected direction — animals in better conditions responded more optimistically — but that the overall effect size was small, and that effect sizes were highly heterogeneous between studies (Lagisz et al., 2020).
That is a genuine validation and a genuine limitation in the same sentence. The paradigm detects something. It detects it weakly, and inconsistently across designs. A parallel meta-analysis of pharmacological manipulations reached a similarly qualified conclusion (Neville et al., 2020).
6.2 The Scoring Method Changes the Answer
The canine-specific methodological analysis is the one practitioners should know about. Examining 51 shelter and 40 pet dogs, the shelter–pet difference appeared when raw approach latencies were analyzed and disappeared under the adjusted score that most published studies use (Burani et al., 2020).
A result that depends on which of two standard analyses is applied is not a robust result. It does not mean the difference is imaginary; it means the measurement is not yet stable enough to carry the weight often placed on it.
6.3 Dogs Drop Out of the Task
The same analysis flagged a further problem: a non-negligible proportion of dogs never completed the training phase and were excluded (Burani et al., 2020). If the dogs that fail to learn a food-location discrimination differ systematically from those that succeed — in motivation, in food value, in anxiety, in age — then the tested sample is not the sample of interest (and individual differences here are substantial).
6.4 The Training Phase Leaves a Mark
The measure is meant to reflect affective state at the time of testing. Krahn et al. (2024) showed that it also reflects what happened before it: dogs given additional discrimination training before the test performed differently in the judgment bias task afterwards.
That does not invalidate the paradigm, and it does mean a latency difference between two groups is not automatically a mood difference — it can also be a difference in how the discrimination was learned. Where groups differ in training history as well as in the variable of interest, and in applied studies they usually do, the two are confounded ([L:/research/operationalizing-dog-behavior-scientific-measurement|which is what operationalizing a behavioral claim is for]).
6.5 Latency Is Not Only About Mood
Approach speed is affected by hunger, general activity level, how much the dog values the food, prior reinforcement history, and arousal. Studies control what they can. The measure remains indirect, and the inferential chain from a walking speed to a mood state is longer than the phrase "optimistic dog" suggests.
6.6 The Small Effect Is the Central Limitation
A real but small cross-species effect (Lagisz et al., 2020; Neville et al., 2020) means group differences are detectable with adequate samples and individual scores are not interpretable. That single fact rules out most of what people want to use the method for.
It is also why the method has not become a clinical tool despite twenty years of interest.
6.7 Training Requirements Limit Who Can Be Tested
The task requires a dog willing to learn two locations and to keep approaching across many trials. Dogs that will not, whether from anxiety, low food motivation or age, are absent from the samples.
That is a selection problem running through the whole literature: the animals most likely to be in a poor affective state are the ones most likely to drop out.
7. How to Read a Judgment Bias Study
7.1 Check Which Probe Produced the Effect
Most designs use three ambiguous positions between the trained ones. An effect at the near-positive probe means something different from an effect at the midpoint, and a study reporting only a combined score has discarded that information.
Where the difference appears only at one probe, the finding is narrower than the headline suggests.
7.2 Check How Dropouts Were Handled
Dogs that stop approaching partway through the session are a known problem, and how a study treats them changes the result. Excluding them removes the animals least willing to engage, which is not a random subset.
A paper that reports its dropout rate is being more honest than one that does not, regardless of what the rate is.
7.3 Check the Scoring Method
Raw latency, latency normalized against the trained positions, and the difference between probe and reference latencies give different answers from the same data, and reported effects can change with the method chosen.
Where two studies disagree, the scoring method is the first thing to compare and it is rarely mentioned in secondary accounts.
7.4 Check What Else Could Move Latency
Speed of approach depends on how fast the dog moves generally, on how motivated it is by the reward on that day, on its age and on any physical discomfort. All of these vary between animals and within an animal across days, which is why a within-animal design is worth more here than a larger between-group sample.
A well-designed study controls for at least some of them. A study that reports only a group difference in latency has measured something, and which something is the open question.
8. What It Means for Practice
8.1 It Is Not a Test You Can Run
The task requires many training trials, controlled conditions, careful timing, and ambiguous probes presented rarely enough that the dog does not learn what they mean. Improvised versions measure nothing. This is a research instrument, and treating it as a practitioner tool misrepresents both.
8.2 What Is Genuinely Usable
Three things transfer. First, the concept: a dog's interpretation of ambiguous situations is shifted by its background emotional state, which reframes the "he's being difficult in new places" complaint as information about mood rather than obedience (the layer beneath trained behavior).
Second, the findings themselves, which give evidence-based weight to positions that otherwise rest on preference: aversive methods and unresolved separation distress are associated with a measurably worse affective state.
Third, the treatment standard. The fluoxetine result establishes that it is possible in principle to ask whether an intervention improved the dog or only quieted it — the question worth asking after any behavior program (including desensitization and counterconditioning).
8.3 What Not to Claim
Not that a specific dog is an optimist or a pessimist. Not that a single observation reveals mood. And not that the absence of a pessimistic response demonstrates a dog is happy — given the small effect sizes and the absence of any validated positive-emotion indicator, that inference runs well past the evidence (as the anxiety literature also shows).
The wording that survives scrutiny is narrower than the wording that circulates. Not "dogs trained with aversives are sad," but "dogs whose owners reported using two or more aversive methods evaluated ambiguous situations more cautiously in a controlled task." The second sentence is less quotable and is the one the data supports.
8.4 What a Practitioner Can Take From It
Three things transfer without any apparatus. That affective state influences how a dog responds to ambiguity, which explains why an anxious dog treats a novel object as a threat and a comfortable one investigates it. That pain shifts this measure, which supports investigating the body first. And that a stranger's stress registers, which is a variable worth managing in a consultation.
None of those requires running the test, and all of them follow from what it has found.
8.5 What to Say When Someone Offers the Test
Commercial claims to assess a dog's optimism are not supported by this literature. The measure produces group comparisons under controlled conditions; an individual score has no established interpretation.
That is not a reason to dismiss the research, which is careful and useful. It is a reason to be specific about what the research supports.
9. What Would Make the Test Usable
9.1 A Shorter Protocol
Training a dog to criterion on two spatial locations takes time that neither a clinic nor a shelter has. A version that produced a usable measure in a single session would change what the method is for.
Nobody has produced one, and the training phase is what makes the measure interpretable, so shortening it is not a simple matter of impatience. Without the trained reference positions there is nothing for the ambiguous one to be ambiguous between.
9.2 Agreement on Scoring
A standard scoring convention, applied across laboratories, would make studies comparable and would resolve part of the current disagreement without any new data being collected.
That is a coordination problem rather than a scientific one, and coordination problems in small fields tend to persist. The multi-lab model that has appeared elsewhere in canine science would solve it, and nobody has applied it here.
9.3 Within-Animal Designs
The most informative version compares the same dog before and after an intervention rather than comparing groups. Individual variation in this task is large enough that between-group comparisons need numbers canine research rarely reaches.
Several of the stronger canine findings already use this design, and it is what makes them stronger. The treatment study discussed earlier is the clearest example in this article.
9.4 Reporting Individuals
A group shift toward faster probe latencies is compatible with most dogs unchanged and a few moving a great deal. Reporting the distribution alongside the mean would tell readers which they are looking at.
This costs nothing and would improve every study in the area immediately.
9.5 Independent Replication
Much of the canine work comes from a small number of connected groups using related protocols. Replication by an unconnected laboratory, using a published protocol in a different population, would do more for confidence in the method than another new finding.
That is true of most of the topics in this collection and is more consequential here, because the method is being proposed as a welfare measurement tool.
10. Which Findings Come From Which Species
10.1 The Paradigm Is Borrowed and the Application Is Canine
The judgment bias approach was developed in rats and formalized as a framework for animal welfare (Harding, Paul & Mendl, 2004; Mendl, Burman & Paul, 2010b), and the cross-species meta-analytic work covers many taxa (Lagisz et al., 2020; Neville et al., 2020).
Everything applying it to dogs is canine, which makes this one of the better-served topics in the collection.
10.2 What Has Been Measured in Dogs
Separation-related behavior (Mendl et al., 2010a; Karagiannis, Burman & Mills, 2015), training method (Casey et al., 2021), chronic pain and neurological conditions (Cockburn et al., 2018; Hobbs et al., 2020), housing and age (Burani et al., 2020; Piotti et al., 2018), the odor of an unfamiliar stressed human (Parr-Cortes et al., 2024), and the effect of brief owner absence (Müller et al., 2012).
That is ten canine studies across seven distinct questions, which is a substantial body of work for a method barely twenty years old.
10.3 What the Cross-Species Work Adds
The meta-analyses are the reason this article can say the effect is real and small (Lagisz et al., 2020; Neville et al., 2020). A single-species literature could not establish that, because a small effect looks like an inconsistent one from inside any one field.
They also supply the correction that matters most: the method detects affective differences at group level and is not a diagnostic instrument for an individual.
10.4 Where the Canine Work Is Weakest
Sample sizes are modest, scoring conventions differ, dropout handling varies, and independent replication of specific canine findings is largely absent. None of that is unusual for the field and all of it limits how confidently a single result can be quoted.
10.5 Why the Method Still Matters
It is currently among the few canine measures that address valence more directly than inferring it from arousal, which is the gap the physiological measures cannot close.
A flawed measure of the right quantity is more useful than a precise measure of the wrong one, provided its limitations are stated.
11. Summary at a Glance
The problem it solves — There is no validated indicator of positive emotion in dogs (Csoltova & Mehinagic, 2020), and standard indicators track arousal rather than valence under positive conditions (Flint et al., 2024).
The logic — Mood biases how ambiguity is interpreted. Trained in rats first (Harding et al., 2004), formalized as a general framework for animal affect (Mendl et al., 2009; 2010b).
The canine task — A bowl in a rewarded location, a bowl in an unrewarded location, and ambiguous positions between. Approach latency is the measure.
What it found — Pessimistic responding in dogs with separation-related behavior (Mendl et al., 2010a), in dogs trained with two or more aversive methods (Casey et al., 2021), in dogs with syringomyelia (Cockburn et al., 2018) and epilepsy (Hobbs et al., 2020).
What shifts it — Treatment with fluoxetine plus behavior modification moved dogs toward less pessimistic responding (Karagiannis et al., 2015), and the odor of an unfamiliar stressed person moved them the other way (Parr-Cortes et al., 2024).
Where it fails — Brief owner absence produced no effect (Müller et al., 2012); the shelter–pet difference appeared under one scoring method and vanished under another (Burani et al., 2020).
The overall effect is small — Across species, manipulations of affect shift judgment bias in the predicted direction, with small and highly heterogeneous effect sizes (Lagisz et al., 2020).
12. Research Gaps and Critical Appraisal
The effect is real but small. The cross-species meta-analysis supports the paradigm's basic validity while showing that effect sizes are small and heterogeneity is high (Lagisz et al., 2020). Any account that presents the test as a reliable mood-reading device overstates what the aggregate data supports.
Analytical choices change conclusions. A canine population difference that appears under raw latencies and disappears under the standard adjusted score (Burani et al., 2020) indicates the measurement has not been methodologically settled.
Sample attrition is unquantified in most reports. Dogs that fail the training phase are excluded, and the ways in which they differ from included dogs are rarely characterized. This is a selection problem sitting underneath a welfare measure.
The valence interpretation is inferred, not demonstrated. Slow approach to an ambiguous bowl is consistent with negative mood. It is also consistent with lower food motivation, higher general caution, or a different reinforcement history. Studies control for some of this; none can eliminate it.
Positive states remain the hard case. Most canine findings concern negative states. Demonstrating that a dog has become positively content — as opposed to less distressed — remains the unsolved problem the paradigm was partly invented to address.
Canine sample sizes are small and mostly cross-sectional. Individual studies typically involve tens of dogs, with limited longitudinal work. The fluoxetine result is the notable exception in design, and it is one study without an untreated control group.
Prior learning contributes to the score. Pre-session discrimination training alters subsequent judgment bias performance in dogs (Krahn et al., 2024), so the measure reflects learning history alongside affective state.
Test–retest reliability in dogs is underreported. Whether a given dog produces a similar bias score on separate occasions is a basic psychometric question that the canine literature has not adequately answered, and without it the measure cannot support individual-level claims.
The effect is real and small. Cross-species meta-analytic work establishes both (Lagisz et al., 2020; Neville et al., 2020), which permits group comparisons and rules out individual diagnosis.
Scoring conventions differ between studies. Raw latency, normalized latency and difference scores produce different answers from the same data, and the choice is rarely discussed in secondary accounts.
Dropout is a selection problem. Dogs that stop approaching are excluded from analysis, and they are plausibly the animals whose affective state the method is most intended to detect.
Independent replication is largely absent. Most canine findings come from a small number of connected groups, so agreement between them reflects shared methodology as much as confirmation.
13. Conclusion
The cognitive bias test is the most serious attempt yet to answer a question that behavior alone cannot settle — whether a dog is in a good or a bad state — and its canine results have been genuinely useful. It supplied evidence that separation-related behavior sits on top of a negative emotional state rather than being a nuisance to be managed, that dogs trained with multiple aversive methods differ measurably in affect and not only in observable stress signals, and that treatment was associated with a shift in the measured affective state rather than only with a quieter dog. It also demonstrated that a stranger's stress odor is enough to move a dog's judgment, which is worth sitting with. At the same time, the aggregate effect across species is small, effect sizes vary widely between designs, a canine population difference can appear and disappear depending on which of two standard scoring methods is applied, and an unquantified proportion of dogs never gets far enough into the task to be tested at all. Both halves are true, and the useful stance holds them together: this is one of the best available instruments the field has for a question it cannot otherwise ask, and it is not yet precise enough to tell you about the dog in front of you. The concept is what practitioners should carry — that background emotional state shapes how a dog reads an ambiguous situation, which reframes a great deal of behavior that otherwise gets attributed to stubbornness.
Key Insights (Takeaways)
The test exists because behavior alone cannot distinguish a calm dog from a shut-down one. There is no validated indicator of positive emotion in dogs (Csoltova & Mehinagic, 2020), and standard indicators track arousal rather than valence under positive conditions (Flint et al., 2024). Judgment bias measures how a dog interprets an ambiguous cue and uses that as a proxy for mood.
The canine findings converge on welfare. Pessimistic responding appeared in dogs with separation-related behavior (Mendl et al., 2010a), in dogs whose owners used two or more aversive training methods (Casey et al., 2021), and in dogs with painful or neurological conditions (Cockburn et al., 2018; Hobbs et al., 2020). Nothing in the canine literature contradicts the welfare framework.
Treatment was associated with a shift in the measured affective state, not only in the behavior. Dogs with separation-related problems treated with fluoxetine plus behavior modification shifted toward less pessimistic responding (Karagiannis et al., 2015) — the first demonstration in a non-human species that clinical treatment changes cognitive bias, and the model for asking whether a dog is better or merely quieter.
A stranger's stress is enough to shift it. Dogs exposed to body odor from an unfamiliar stressed person responded more pessimistically than those exposed to odor from a relaxed person (Parr-Cortes et al., 2024). The transfer of human emotional state to dogs is not limited to visible cues or familiar people.
The method is weaker than the headlines. Across species the effect is small and highly heterogeneous (Lagisz et al., 2020); in dogs, a shelter–pet difference appeared under raw latencies and disappeared under the standard adjusted score, with a meaningful proportion of dogs excluded for failing training (Burani et al., 2020), and the score also reflects prior discrimination training (Krahn et al., 2024). It is a research instrument, not a test to run on an individual dog.
References
Burani, C., Barnard, S., Wells, D., Pelosi, A., & Valsecchi, P. (2020). Using judgment bias test in pet and shelter dogs (Canis familiaris): Methodological and statistical caveats. PLoS ONE, 15(10), e0241344. https://doi.org/10.1371/journal.pone.0241344
Casey, R. A., Naj-Oleari, M., Campbell, S., Mendl, M., & Blackwell, E. J. (2021). Dogs are more pessimistic if their owners use two or more aversive training methods. Scientific Reports, 11(1), 19023. https://doi.org/10.1038/s41598-021-97743-0
Cockburn, A., Smith, M., Rusbridge, C., Fowler, C., Paul, E. S., & Murrell, J. C. (2018). Evidence of negative affective state in Cavalier King Charles Spaniels with syringomyelia. Applied Animal Behaviour Science, 201, 77–84. https://doi.org/10.1016/j.applanim.2017.12.019
Csoltova, E., & Mehinagic, E. (2020). Where do we stand in the domestic dog (Canis familiaris) positive-emotion assessment: A state-of-the-art review and future directions. Frontiers in Psychology, 11, 2131. https://doi.org/10.3389/fpsyg.2020.02131
Flint, H. E., Weller, J. E., Parry-Howells, N., Ellerby, Z. W., McKay, S. L., & King, T. (2024). Evaluation of indicators of acute emotional states in dogs. Scientific Reports, 14(1), 6406. https://doi.org/10.1038/s41598-024-56859-9
Harding, E. J., Paul, E. S., & Mendl, M. (2004). Cognitive bias and affective state. Nature, 427(6972), 312. https://doi.org/10.1038/427312a
Hobbs, S. L., Law, T. H., Volk, H. A., Younis, C., Casey, R. A., & Packer, R. M. A. (2020). Impact of canine epilepsy on judgement and attention biases. Scientific Reports, 10(1), 17719. https://doi.org/10.1038/s41598-020-74777-4
Karagiannis, C. I., Burman, O. H. P., & Mills, D. S. (2015). Dogs with separation-related problems show a "less pessimistic" cognitive bias during treatment with fluoxetine (Reconcile™) and a behaviour modification plan. BMC Veterinary Research, 11, 80. https://doi.org/10.1186/s12917-015-0373-1
Krahn, J., Azadian, A., Cavalli, C., Miller, J., & Protopopova, A. (2024). Effect of pre-session discrimination training on performance in a judgement bias test in dogs. Animal Cognition, 27(1), 66. https://doi.org/10.1007/s10071-024-01905-2
Lagisz, M., Zidar, J., Nakagawa, S., Neville, V., Sorato, E., Paul, E. S., Bateson, M., Mendl, M., & Løvlie, H. (2020). Optimism, pessimism and judgement bias in animals: A systematic review and meta-analysis. Neuroscience & Biobehavioral Reviews, 118, 3–17. https://doi.org/10.1016/j.neubiorev.2020.07.012
Mendl, M., Brooks, J., Basse, C., Burman, O., Paul, E., Blackwell, E., & Casey, R. (2010a). Dogs showing separation-related behaviour exhibit a 'pessimistic' cognitive bias. Current Biology, 20(19), R839–R840. https://doi.org/10.1016/j.cub.2010.08.030
Mendl, M., Burman, O. H. P., Parker, R. M. A., & Paul, E. S. (2009). Cognitive bias as an indicator of animal emotion and welfare: Emerging evidence and underlying mechanisms. Applied Animal Behaviour Science, 118(3–4), 161–181. https://doi.org/10.1016/j.applanim.2009.02.023
Mendl, M., Burman, O. H. P., & Paul, E. S. (2010b). An integrative and functional framework for the study of animal emotion and mood. Proceedings of the Royal Society B: Biological Sciences, 277(1696), 2895–2904. https://doi.org/10.1098/rspb.2010.0303
Müller, C. A., Riemer, S., Rosam, C. M., Schößwender, J., Range, F., & Huber, L. (2012). Brief owner absence does not induce negative judgement bias in pet dogs. Animal Cognition, 15(5), 1031–1035. https://doi.org/10.1007/s10071-012-0526-6
Neville, V., Nakagawa, S., Zidar, J., Paul, E. S., Lagisz, M., Bateson, M., Løvlie, H., & Mendl, M. (2020). Pharmacological manipulations of judgement bias: A systematic review and meta-analysis. Neuroscience & Biobehavioral Reviews, 108, 269–286. https://doi.org/10.1016/j.neubiorev.2019.11.008
Parr-Cortes, Z., Müller, C. T., Talas, L., Mendl, M., Guest, C., & Rooney, N. J. (2024). The odour of an unfamiliar stressed or relaxed person affects dogs' responses to a cognitive bias test. Scientific Reports, 14(1), 15843. https://doi.org/10.1038/s41598-024-66147-1
Piotti, P., Szabó, D., Bognár, Z., Egerer, A., Hulsbosch, P., Carson, R. S., & Kubinyi, E. (2018). Effect of age on discrimination learning, reversal learning, and cognitive bias in family dogs. Learning & Behavior, 46(4), 537–553. https://doi.org/10.3758/s13420-018-0357-7