Behavioral Assessment in Dogs: From Description to Plan
Michael Sauerwein · September 25, 2026
A behavior assessment has to answer four questions in order: what exactly does the dog do, in which situations, what factors may maintain it, and what else could explain it. Most disagreements about cases come from skipping one of them.
This article sets out what the canine assessment instruments actually deliver, why description has to come before interpretation, how function is determined rather than guessed, which differentials belong in every case, and how to make a plan measurable enough to be evaluated.
1. What an Assessment Has to Do
1.1 Four Questions, Not One
A behavior assessment has to answer four questions in order: what exactly does the dog do, in which situations does it happen, what is maintaining it, and what else could explain it. Most disagreements about cases come from skipping one of them — usually the first or the last.
Everything that follows is an attempt to answer those four questions with methods that another person could repeat and check.
1.2 Description Before Interpretation
"He's dominant", "she's anxious", "he's protecting me" are conclusions. "He barks and lunges to the end of the lead when a dog appears within about fifteen metres, and takes two or three minutes to take food again afterwards" is an observation. Only the second can be tested, tracked over time, or handed to a colleague (how behavior is defined before it is measured).
1.3 How to Read This Article
The methodological findings come from systematic reviews of the canine assessment literature. The clinical sequence draws on veterinary behavioral medicine, and the functional analysis section on the small applied literature that has tested those methods in dogs. Each section says which.
1.4 Why Labels Travel Badly
Labels are efficient: one word carries a whole picture, and everyone in the room nods. The problem appears at the next step. Two trainers who both call a dog "reactive" may be describing a dog that barks at other dogs from twenty metres and a dog that redirects onto the lead, and a plan built for one will fail for the other.
Descriptions travel. A record that says what the dog did, at what distance, and how long it took to recover means the same thing to the next person who reads it, including the veterinarian.
1.5 Assessment Is Not a Test
Households often expect an assessment to be an event: bring the dog, observe it, receive a verdict. In reality, no single observation can reveal a dog’s general character or reliably predict future behavior. What actually produces useful information is a process spread over time — history, records from home, video, observation in the situations that matter, and revision as the plan runs.
Saying that in the first conversation prevents a common disappointment and sets up the household to collect what the assessment needs.
1.6 Who Does the Assessment
In practice the work is split. Owners supply the history and the everyday records, trainers and behavior consultants supply structured observation and the functional analysis, and veterinarians supply the medical differentials and any diagnosis. Problems arise when one part of that chain claims the whole of it, in either direction.
The chain also determines what a written assessment is for: it has to be readable by the next person in it, which is a stronger requirement than being useful to its author.
2. The State of the Instruments
2.1 Questionnaires
Questionnaires are the most widely used assessment tool in canine research and practice. A systematic review examined how they characterize behavior: 38 studies reporting 37 questionnaires, of which 13 were versions of the validated C-BARQ, 13 were unpublished or original instruments and 11 were other validated questionnaires (Johnson, Schenke, Foster & Stephens-Lewis, 2026).
Item coding revealed substantial variability in how behavior was operationalized: emotion-based descriptors ranged from 2.3 to 40.5 percent across groups, and 26.3 to 64.4 percent of items were mixed or could not be categorized, indicating that observable behavior is routinely conflated with inferred traits or emotional states (Johnson et al., 2026).
The authors also report that their planned assessment of reliability and validity could not be carried out, because reporting in the source papers was insufficient and often limited to subscales.
2.2 What That Means for Owner Questionnaires
An owner answering "how anxious is your dog around strangers" is being asked to judge an internal state. An owner answering "how often does your dog bark when a stranger enters the house" is being asked what happened. The first is quicker and the second is checkable, and most instruments mix the two without saying so.
The C-BARQ remains the most widely validated example in the field, and it was developed specifically to measure behavior and temperament traits through owner reports of behavior in defined contexts (Hsu & Serpell, 2003).
2.3 Behavioral Tests
A scoping review of behavioral testing analyzed 392 publications from 1948 to 2024, reporting 2,362 individual tests. Indices of test reliability were reported in 39.5 percent of articles and indices of validity in 32.9 percent, with construct validity in 20.2 percent and criterion validity in 14.5 percent (Moser, Welch, Brown, McGreevy & Bennett, 2024).
In other words, most published canine behavioral tests do not report whether they measure what they claim or whether two observers would score them alike.
2.4 Working-Dog Selection Tests
A systematic review of behavioral tests used to assess characteristics important in working dogs approached the same problem from the applied side, motivated by the number of dogs withdrawn from service despite passing selection (Brady, Cracknell, Zulch & Mills, 2018).
That combination — high stakes, widespread testing, uncertain predictive value — recurs wherever tests are used to make decisions about individual dogs.
2.5 Direct Observation
A systematic review of observational welfare measures included 39 studies and identified nine overarching themes of behavioral indicators, most commonly vocalizations, stress-related behaviors and interaction with the non-social environment. Only five of the articles mentioned any form of validity assessment, 23 reported inter-rater reliability, and only six linked their observations to a concrete conclusion about the dogs' welfare (De Winkel et al., 2024).
2.6 The Honest Summary
Across questionnaires, tests and direct observation, the field measures a great deal and validates little. That is not an argument for abandoning assessment; it is an argument for treating any instrument's output as one piece of evidence rather than as a verdict.
2.7 Why Inference Creeps In
Asking people to report only behavior is harder than it sounds. "Was your dog frightened?" is easy to answer; "what did your dog do with its body and where did it move" requires attention and memory. Instruments take the easy route because response rates depend on it, and the result is the conflation the review documents (Johnson et al., 2026).
The same pressure operates in a consultation. The useful counter-move is to ask for one concrete recent episode in detail rather than a general characterization.
2.8 Tests Measure the Situation Too
A behavioral test result reflects the dog, the test, the tester and the day. It is a sample of behavior under specific conditions, not a complete measurement of the dog. The scoping review documents just how varied those testing procedures are: 2,362 individual tests across 392 publications, with limited standardization (Moser et al., 2024).
Where a test is used to make a decision about an individual dog, its predictive value has to come from evidence rather than from the test's face plausibility — the point the working-dog review makes from the applied side (Brady et al., 2018).
2.9 Reliability Before Validity
Two questions come before any interpretation: would a second observer score this the same way, and would the same dog score similarly next week? Inter-rater reliability was reported in 23 of 39 observational welfare studies (De Winkel et al., 2024) and in about a third of behavioral testing publications (Moser et al., 2024), which is more often than validity and still not routine.
2.10 Where Questionnaires Are Genuinely Useful
Despite the limitations, standardized questionnaires do things a conversation cannot: they cover domains the household would not have mentioned, they allow comparison with a reference population, and they produce a record that can be repeated later to track change. The C-BARQ's value lies precisely in being administered the same way to many dogs (Hsu & Serpell, 2003).
The honest use is as a screening and tracking instrument rather than as a diagnosis, and with awareness that some of its items ask for inference rather than observation (Johnson et al., 2026).
2.11 Why Standardization Is Hard Here
Dogs are assessed in the situations that matter to them, and those situations are not standardizable across households. A test that is identical everywhere loses the context; an assessment in context loses comparability. The field's methodological variability is partly a consequence of that trade-off rather than mere carelessness (Moser et al., 2024).
3. The History
3.1 What to Collect
A usable history covers: what the dog does, described in observable terms; when it started; what was happening in the household at that time; how often it occurs now and whether that has changed; what the situation looks like immediately before and after; who is present; what the people do; and what has been tried.
3.2 The Timeline Does Most of the Work
Onset and course separate more cases than any other single piece of information. A gradual change with no external cause points differently from a sudden change after a specific event, and a problem that escalates in steps points differently from one that has been stable for years (why the timeline matters medically).
3.3 Video Is Not Optional
Home video gives what a consultation cannot: the behavior in its own context, the sequence before it, the household's actual responses rather than their description of them, and the dog's movement and rest. Structured approaches in veterinary behavioral medicine use home video analysis as a core component alongside history and clinical observation (Kwik, De Keuster, Bosmans & Mottet, 2025).
3.4 The Owner's Account Is Evidence and Interpretation
Owners are the only source of long-term information about the dog, and their account arrives pre-interpreted: what they noticed, what they thought it meant, what they remember because it fit. The task is not to distrust it but to ask for the observations underneath the conclusions.
3.5 What Did Not Happen Is Information
Situations in which the behavior does not occur are as informative as those in which it does. A dog that lunges at dogs on the street but not at the same dogs off-lead in a field, or that growls at one household member and not another, has narrowed the problem considerably before any plan is written.
Households rarely volunteer this, because they are reporting a problem rather than mapping it. It has to be asked for directly.
3.6 Sequence, Not Snapshot
Most useful information sits in the seconds around the behavior: what the dog was doing before, the first sign it gave, what the person did, and what happened next. Recording that sequence turns "he bites without warning" into a description that almost always contains a warning nobody had noticed (how reactive episodes unfold).
3.7 What the Household Has Already Tried
Previous attempts carry two kinds of information: what has already been ruled out, and what the dog has learned in the meantime. A dog that has been punished for growling has a different learning history from one that has not, and that changes both the prognosis and the plan (what punishment adds to a case).
3.8 Household Agreement
Different people in the same home often respond differently to the same behavior, and the resulting inconsistency shapes what the dog has learned. Asking each person separately what they do produces a more accurate picture than asking the household collectively, where one account usually dominates (why inconsistency builds persistence).
3.9 What the Dog Cannot Avoid
Between appointments the dog continues to meet the triggering situations, and those unplanned encounters are part of the case. An assessment that does not map them — the stairwell, the school run past the fence, the neighbor's dog at 6 pm — will produce a plan that is undone daily.
3.10 Rest, Routine and the Rest of the Day
The behavior occupies minutes; the day it sits in lasts sixteen hours. How much uninterrupted sleep the dog gets, how predictable the routine is, how much of the day it spends alone, and what happens in the hours before the problem all belong in the history because they change the baseline the behavior occurs against (how canine rest is organized).
4. Determining Function
4.1 The Question
Function asks what the behavior produces: attention, access to something, the end of something unpleasant, or a sensory consequence. Two dogs performing identical behaviors can be maintained by different consequences, and the plans differ accordingly.
4.2 Functional Analysis in Dogs
Functional analysis — testing candidate consequences experimentally rather than inferring them — has been applied to dogs. Access to a familiar person can function as a reinforcer in a functional analysis format (Feuerbacher & Wynne, 2016), and the approach has been used with problem behaviors including jumping up (Pfaller-Sadovsky, Arnott & Hurtado-Parrado, 2019).
4.3 The Practical Version
Full functional analysis is rarely feasible in a household. The workable substitute is systematic observation of what reliably follows the behavior, plus one low-risk test: withhold the suspected consequence once, in a safe situation, and see whether the behavior changes.
That is weaker than an experiment and far stronger than assuming the function from the behavior's appearance (how function shapes a plan).
4.4 Emotion and Function Are Different Questions
Asking what a behavior produces does not answer what the dog feels, and some presentations are driven by distress rather than by an expected consequence. Treating a fear-driven behavior as an operant problem is the most consequential version of getting this wrong (why the mechanism decides the plan).
4.5 Multiple Functions at Once
A behavior can serve more than one function, and the functions can change over time. A dog that first barked at visitors out of fear may now also be barking because it has reliably produced retreat. Where both are present, addressing only one leaves the behavior standing.
4.6 Testing a Hypothesis Safely
Any test of function has to be safe for everyone involved, which rules out testing aggression by provoking it. In those cases the function is inferred from careful observation and from what changes when management removes the situation, rather than from an experiment (why arranging the situation matters).
4.7 Sensory and Attentional Explanations
Some apparent behavior changes are perceptual: a dog with reduced hearing startles when approached from behind, a dog with failing sight reacts to shapes at dusk. Those belong in the same layer as pain — physical explanations that produce behavioral presentations and are invisible unless someone asks (what is known about canine hearing).
4.8 Emotional State Cannot Be Measured Directly
Assessments regularly record "the dog was anxious". What was observed is behavior; the anxiety is an inference, and it may be right. Keeping the two visibly separate in the record — observation in one column, interpretation in another — costs nothing and makes the reasoning checkable (why behavior is not an identified state).
4.9 Motivation Changes With Context
A dog that will work for food in the kitchen may refuse it near the trigger, and that refusal is data rather than stubbornness: it usually marks the point at which the situation has become too much. Using food acceptance as an indicator in the assessment turns a vague judgment about stress into an observable threshold (how arousal shows itself).
5. Differential Diagnosis
5.1 The Medical Layer
Pain and other medical conditions belong on the differential for aggression, fear and phobia responses, sleep disturbance, inappropriate elimination, vocalization and repetitive behaviors, because the most common sign of pain in animals is a change in behavior (Camps, Amat & Manteca, 2019).
Estimates of medical involvement in behavior caseloads have risen over the decades, from around 5 percent in early reports to roughly a quarter in a referral caseload, and the authors of the review describe that as at the lower end of what has been reported more recently (Mills et al., 2020).
5.2 What Cannot Be Excluded From the Behavior Alone
No behavioral sign identifies pain. Signs are often non-specific and may be absent during a clinical examination (Kwik et al., 2025), which is why the assessment's job is to recognize when a workup is warranted rather than to rule pain in or out.
5.3 The Developmental Layer
Age changes what is expected. Some presentations in young dogs are developmentally ordinary, and behavior profiles shift as dogs mature (what changes in adolescence). At the other end, new behavior in an older dog raises different questions (what changes in aging dogs).
5.4 The Environmental Layer
Rest, routine, activity and what the household has been doing differently all belong in the assessment. A dog that cannot settle may be under-rested rather than under-exercised, and the two require opposite responses (why more activity is not always the answer).
5.5 The Learning Layer
Whatever else is going on, a behavior that has occurred repeatedly has a learning history. Pain, fear and learning are not competing explanations so much as layers, and treating one usually leaves the others in place (why learned behavior persists).
That is why "the pain is treated, so the behavior should stop" disappoints so often. The dog that learned that growling ends the handling has learned something that outlives the joint.
5.6 Working Order of the Differentials
A workable order: exclude or address medical causes where anything changed; check rest, routine and basic needs; establish whether the presentation is distress-driven; and only then treat it as primarily a learning problem. The order reflects the cost of being wrong rather than the likelihood of each explanation.
5.7 Sensitization and Habituation as Explanations
A problem that has grown gradually under repeated exposure may not need a new cause at all: repeated exposure above a dog's threshold increases responding rather than reducing it. Recognizing that pattern in the history often explains a case that otherwise looks like deterioration for no reason (how repetition changes responding).
5.8 Behavioral Medication Is a Veterinary Decision
Where an assessment points toward a condition that may warrant medication, the decision belongs to a veterinarian, and the assessment's contribution is the description and the baseline measures that allow the effect to be judged (how behavioral medication is used).
That is also the strongest practical argument for measurable outcomes: without them, a medication trial is evaluated on impressions formed by people who hope it works.
5.9 Two Errors With Different Costs
Treating a medical case as a training problem costs the dog time in an unresolved condition. Treating a training problem as medical costs a consultation. That asymmetry is the whole argument for keeping the medical differential open early rather than turning to it after a plan has failed (why pain is easy to miss).
5.10 Breed Expectations Are Not a Differential
"It's a herding breed, that's why" explains nothing about why this dog does this now. Breed shifts probabilities and explains a modest share of behavioral variation between individuals; it does not identify what maintains a behavior or what else could be causing it (how much breed explains).
Where breed is genuinely useful is in setting expectations about what a dog finds reinforcing, which feeds into the plan rather than into the diagnosis.
6. Making It Measurable
6.1 Choosing the Measure
A usable outcome measure is something a household can record without training: frequency, duration, latency to settle, the distance at which the dog first reacts, or the intensity on a simple scale defined in advance.
6.2 Baseline First
Without a baseline, improvement is an impression. Two weeks of simple records before the plan starts is usually enough, and it also reveals patterns nobody had noticed — the days, times and people that cluster with the behavior.
6.3 Define Success in Advance
Agreeing what would count as progress prevents two failures: abandoning a plan that is working slowly, and continuing one that is not working at all. Realistic markers are fewer episodes, shorter episodes, lower intensity and faster recovery.
6.4 Physiological Measures Are Not the Answer Here
Cortisol and heart rate variability measure activation rather than how a dog evaluates its situation, and they vary with sampling, individual and context (what stress measures capture). In an applied assessment they add cost and rarely change a decision.
6.5 Who Records What
The household records the everyday measure — a tally, a time, a distance. The professional records the assessment itself, the plan and its criteria. Splitting it that way keeps the burden small enough that the records actually get made, which is the usual reason they do not.
6.6 Video as the Reliability Check
Where a measure matters, a short video of the same situation each week is the cheapest reliability check available: it allows the same behavior to be scored twice, by two people, without relying on memory. That is the practical version of the inter-rater question the reviews raise (De Winkel et al., 2024).
6.7 Beware of Measuring the Easy Thing
Counting barks is easy; whether the dog can settle afterwards may matter more. A measure that is convenient but peripheral produces a record that improves while the problem does not, which is worse than no record at all.
6.8 How Long a Baseline Needs to Be
Long enough to cover the variation. For a behavior that happens daily, a week or two suffices; for one that happens when visitors come, the baseline has to span enough visits to be meaningful, which may be a month. Judging a plan against a baseline that was too short is a common way to conclude the wrong thing.
6.9 Regression to the Mean
Households usually seek help when things are at their worst. Some improvement afterwards would have happened anyway, which is why a pre-plan baseline and a defined criterion matter more than the impression that things are better than at the first appointment.
6.10 Scales Need Anchors
A three-point intensity scale is useful only if each point is described: what counts as a one, a two and a three, in observable terms, written down before the recording starts. Without anchors the scale drifts with the recorder's mood, and a month of data becomes uninterpretable.
Anchors also make the record transferable: two people in the same household can score the same episode and compare, which is the household version of inter-rater reliability.
6.11 Recording Should Survive a Bad Week
The record has to be simple enough that it still happens when the household is tired, busy or discouraged. One line a day beats a detailed form that is abandoned after ten days, and a record with gaps is more useful than no record, provided the gaps are visible rather than filled in from memory.
7. From Description to Plan
7.1 What a Finished Assessment Contains
An operational description of the behavior; the situations in which it occurs and does not occur; a timeline; the likely function or functions; the differentials considered and what was done about them; a baseline measure; and the plan with its success criteria.
7.2 Where Assessment Stops and Treatment Begins
An assessment that produces no testable statement is not finished. "Fear of strangers" is a label; "approaches and retreats repeatedly when a stranger enters, does not take food within two metres, recovers within a minute when the stranger sits down" is a description that implies where to start.
7.3 Reassessment
The first plan is a hypothesis. If the measures do not move, the assessment was wrong somewhere — usually in the function, the differential or the intensity at which the plan operates — and the honest response is to go back rather than to apply the same plan harder.
7.4 A Worked Example
Presenting complaint. "He's aggressive with visitors."
Description. Barks from the moment the doorbell rings, follows the visitor at two to three metres barking intermittently for five to ten minutes, takes food after about four minutes, has snapped once when a visitor reached toward him while he was on the sofa.
Situations. Every visitor, worse with men, not present when the household meets the same people outdoors, absent in the second hour of a visit.
Timeline. Gradual over eight months, no obvious trigger event, no change in the household.
Function. Barking has reliably produced distance — visitors stop moving and the household puts the dog behind a gate. The snap produced immediate retreat.
Differentials. Gradual onset in an adult dog with no external change puts a veterinary workup first; the sofa incident involves handling and proximity to a resting place.
Measure. Time to first food acceptance after the visitor sits down, recorded at each visit.
That is an assessment. The training plan follows from it, and any part of it can be checked by somebody else.
7.5 What to Write Down for a Colleague
A referral note that contains the description, the situations, the timeline and the current measure is useful to a veterinarian, a behavior specialist or the next trainer. One that contains a label and a method is not, because none of it can be verified or built on.
7.6 When the Assessment Says "Refer"
Three findings send a case elsewhere rather than into a training plan: a medical pattern, a presentation that involves serious risk to people, and a picture that does not fit any coherent account after a careful assessment. Recognizing the limits of one's own remit is part of the assessment rather than a failure of it.
7.7 What to Tell a Household
Four sentences cover it. I need to know exactly what your dog does, not what you think it means. I need to know when it does not happen as well as when it does. Some of the possible explanations are medical, so the veterinarian may come before the training. And we will agree in advance what counts as progress, so that we both know whether this is working.
7.8 Revisiting the Assessment
Plans are usually revised for one of four reasons: the function was misread, a differential was missed, the intensity was wrong, or the household could not carry out the plan as written. Checking those four in order is faster than starting again, and the fourth is the one professionals check last and should check first.
7.9 Safety Belongs in the Assessment
Where a case involves biting, children, or a dog that cannot be handled, the assessment includes an explicit statement of what must not happen while the plan runs: which situations are avoided entirely, what equipment is used, and who is responsible for each. That is part of the assessment rather than an appendix to it, because it determines what the plan can realistically attempt.
8. Summary at a Glance
Questionnaires conflate behavior with inferred states — Mixed or uncategorizable items ranged from 26.3 to 64.4 percent across groups in a review of 37 instruments (Johnson et al., 2026).
Most behavioral tests do not report validity — Reliability indices in 39.5 percent and validity indices in 32.9 percent of 392 publications (Moser et al., 2024).
Observational welfare measures are rarely validated — Five of 39 studies mentioned any validity assessment, and six linked observations to a welfare conclusion (De Winkel et al., 2024).
Medical differentials belong in every assessment — Behavior change is the most common sign of pain in animals (Camps et al., 2019).
Function cannot be read off the behavior — Identical behaviors can be maintained by different consequences (Feuerbacher & Wynne, 2016).
A baseline turns impressions into data — And costs a household two weeks of simple records.
8.1 What This Means for Practice
The reviews are not an indictment of applied work; they describe a field whose tools have outrun their validation. The reasonable response is neither to reject the instruments nor to treat their output as fact, but to use them as one source among history, observation and measured outcomes — and to be explicit about which claims rest on which.
8.2 Where This Article Sits in the Library
This is the procedural counterpart to several topics covered separately: defining behavior before measuring it, determining function, recognizing medical contributions, and reading physiological measures with appropriate caution. The assessment is where those come together in one case, in one order, for one dog.
8.3 Why Assessment Quality Decides Outcomes
Most published debate in dog training is about methods, and most failures in practice are about assessment: the wrong function, a missed medical cause, an intensity the dog cannot handle, or a plan the household cannot run. Methods differ in their costs and in what they teach the dog; a plan aimed at the wrong problem fails regardless of which method it uses.
That is the case for spending the first appointment on questions rather than on exercises, and for writing the answers down.
9. Research Gaps and Critical Appraisal
Reporting standards are the bottleneck. A planned assessment of reliability and validity could not be completed because the source papers reported too little (Johnson et al., 2026).
No standard assessment protocol exists. Veterinary behavioral medicine and applied behavior analysis each have their own sequence, and no comparison of outcomes between approaches has been published.
Functional analysis in dogs rests on small studies. Single-case designs demonstrate feasibility rather than population-level performance (Feuerbacher & Wynne, 2016; Pfaller-Sadovsky et al., 2019).
Prevalence of medical contribution is uncertain. The figures come from referral caseloads (Mills et al., 2020).
Owner-recorded measures are unvalidated. The practical measures recommended here have not been tested for reliability against observer coding.
9.1 What Would Move the Field
Three things. Reporting standards for reliability and validity in canine behavioral instruments, so reviews can do the synthesis they currently cannot (Johnson et al., 2026). Validation of the simple owner-recorded measures that applied work actually relies on. And outcome studies comparing assessment approaches, so that the sequence recommended here can be tested rather than argued for.
9.2 A Note on Scope
Nothing here describes veterinary diagnosis, and the sequence is not a substitute for a behavioral medicine consultation in cases involving aggression toward people, severe distress or suspected medical causes. What it describes is the part that any competent practitioner can do well: observe precisely, collect a history, form a testable hypothesis about function, keep the differentials visible, and measure whether the plan is working.
9.3 Reading a New Assessment Tool
When a new questionnaire, test or app appears, four questions settle most of it. What does it claim to measure, and is that a behavior or an inferred state? Has anyone reported whether two raters agree? Has it been checked against anything external? And what decision is it supposed to inform — because an instrument that is adequate for research comparisons may be far too coarse for a decision about one dog.
10. Conclusion
Behavioral assessment in dogs is better described as a sequence than as a test. It starts with an operational description, because the instruments in widest use mix observable behavior with inferred traits and emotional states — in a review of 37 canine questionnaires, between 26.3 and 64.4 percent of items could not be categorized cleanly (Johnson et al., 2026). It relies on history and home video rather than on a single consultation, because signs are frequently absent when someone is watching (Kwik et al., 2025). It asks what the behavior produces rather than what it looks like, a question that has been answered experimentally in dogs in a small applied literature (Feuerbacher & Wynne, 2016). It carries a medical differential throughout, since behavior change is the most common sign of pain (Camps et al., 2019) and estimates of medical involvement in behavior caseloads have risen as clinicians have looked harder (Mills et al., 2020). And it ends with something measurable, because the field's own reviews show how rarely its tools report whether they measure what they claim (Moser et al., 2024; De Winkel et al., 2024). None of that requires equipment. It requires writing down what the dog does, when, and what happens next — and being willing to find that the first hypothesis was wrong.
Key Insights (Takeaways)
Assessment answers four questions in order: what, when, what maintains it, and what else could explain it.
Widely used questionnaires mix observable behavior with inferred states, which limits what their scores mean (Johnson et al., 2026).
Most canine behavioral tests do not report reliability or validity (Moser et al., 2024).
History, timeline and home video carry more weight than a single consultation (Kwik et al., 2025).
Function has to be determined rather than inferred from what the behavior looks like (Feuerbacher & Wynne, 2016).
Medical differentials belong in every case, not only the difficult ones (Camps et al., 2019; Mills et al., 2020).
A baseline and a pre-agreed success criterion decide whether a plan can be evaluated at all.
References
Brady, K., Cracknell, N., Zulch, H., & Mills, D. S. (2018). A systematic review of the reliability and validity of behavioural tests used to assess behavioural characteristics important in working dogs. Frontiers in Veterinary Science, 5, 103. https://doi.org/10.3389/fvets.2018.00103
Camps, T., Amat, M., & Manteca, X. (2019). A review of medical conditions and behavioral problems in dogs and cats. Animals, 9(12), 1133. https://doi.org/10.3390/ani9121133
De Winkel, T., van der Steen, S., Enders-Slegers, M.-J., Griffioen, R., Haverbeke, A., Groenewoud, D., & Hediger, K. (2024). Observational behaviors and emotions to assess welfare of dogs: A systematic review. Journal of Veterinary Behavior, 72, 1–17. https://doi.org/10.1016/j.jveb.2023.12.007
Feuerbacher, E. N., & Wynne, C. D. L. (2016). Application of functional analysis methods to assess human-dog interactions. Journal of Applied Behavior Analysis, 49(4), 970–974. https://doi.org/10.1002/jaba.318
Hsu, Y., & Serpell, J. A. (2003). Development and validation of a questionnaire for measuring behavior and temperament traits in pet dogs. Journal of the American Veterinary Medical Association, 223(9), 1293–1300. https://doi.org/10.2460/javma.2003.223.1293
Johnson, A., Schenke, K., Foster, J., & Stephens-Lewis, D. (2026). A systematic review of the characterization of behavior in canine behavioral questionnaires. Journal of Applied Animal Welfare Science. Advance online publication. https://doi.org/10.1080/10888705.2026.2693645
Kwik, J., De Keuster, T., Bosmans, T., & Mottet, J. (2025). Detection of maladaptive pain in dogs referred for behavioral complaints: Challenges and opportunities. Frontiers in Behavioral Neuroscience, 19, 1569351. https://doi.org/10.3389/fnbeh.2025.1569351
Mills, D. S., Demontigny-Bédard, I., Gruen, M., Klinck, M. P., McPeake, K. J., Barcelos, A. M., Hewison, L., Van Haevermaet, H., Denenberg, S., Hauser, H., Koch, C., Ballantyne, K., Wilson, C., Mathkari, C. V., Pounder, J., Garcia, E., Darder, P., Fatjó, J., & Levine, E. (2020). Pain and problem behavior in cats and dogs. Animals, 10(2), 318. https://doi.org/10.3390/ani10020318
Moser, A. Y., Welch, M., Brown, W. Y., McGreevy, P., & Bennett, P. C. (2024). Methods of behavioral testing in dogs: A scoping review and analysis of test stimuli. Frontiers in Veterinary Science, 11, 1455574. https://doi.org/10.3389/fvets.2024.1455574
Pfaller-Sadovsky, N., Arnott, G., & Hurtado-Parrado, C. (2019). Using principles from applied behaviour analysis to address an undesired behaviour: Functional analysis and treatment of jumping up in companion dogs. Animals, 9(11), 1091. https://doi.org/10.3390/ani9111091