Zum Inhalt springen
unterHUNDs – Hundeschule und Verhaltenstherapie im Saarland Initiative für gewaltfreies Hundetraining

Research

Puppy Temperament Tests: What They Predict and What They Do Not

Michael Sauerwein · September 8, 2026

A young puppy sitting calmly on the floor while a person observes and makes notes on a clipboard, illustrating a behavioral assessment session.

A breeder runs a litter through a series of exercises at seven weeks — a startle noise, a restraint hold, a retrieve, a following test — and assigns each puppy a profile. Families are matched accordingly, working prospects are identified, and the results are treated as information about the dog these puppies will become.

Three separate lines of evidence suggest that the predictive information available from such tests is limited. A longitudinal study following Border Collies from the first days of life into adulthood found a single behavior that carried over. A meta-analysis of thirty-one studies found personality consistency in puppies at roughly half the level seen in adults, and lowest for precisely the traits that tests are used to assess. And one of the largest service-dog datasets found the correspondence between puppy test and adult performance negligible — while the puppy test traits themselves turned out to be substantially heritable, which makes the failure more interesting rather than less.

This article covers what the tests measure, why prediction fails at this age, what selection can defensibly use, and why the practice continues regardless (with the same measurement structure as shelter behavioral assessment).

1. What Puppy Tests Claim

1.1 The Basic Proposition

That behavior observed in a puppy at six to eight weeks reflects a stable disposition, and that the disposition will express itself in the adult dog in recognizable form.

From that follows the practical use: match the confident puppy to the active family, the quiet one to the retired couple, and select the one with the right profile for detection work.

1.2 What Gets Measured

The standard batteries assess social attraction, following, restraint tolerance, social dominance, elevation tolerance, retrieving, touch sensitivity, sound sensitivity and sight sensitivity — usually scored on a numerical scale with a composite profile at the end.

The specific items vary between protocols and the structure is remarkably consistent: brief exposures, a single session, a numerical output.

1.3 What Is Being Inferred

The inference chain has three steps, and each requires the previous one. The behavior observed reflects an underlying trait rather than the situation. The trait is stable across development. And the trait expresses in adulthood in a form the test predicts.

Every step is testable, and the evidence bears on all three (with the general inference problem set out separately).

1.4 The Two Uses Are Different Problems

Placement matching and working selection ask different things of the same instrument, and they fail differently.

Matching asks which puppy suits which household — a relative judgment within a litter, where being roughly right may be adequate and where the family adapts to whatever they get.

Selection asks whether this individual will meet an absolute standard eighteen months out, where a wrong answer costs a full training investment.

Selection is the harder problem and the one with the better data, which is why the service-dog literature carries most of the weight in this article.

1.5 Why It Is Plausible

Individual differences in dogs are real, partly heritable and stable in adults. If personality exists, and it does, then measuring it early looks like a reasonable ambition rather than a category error.

The question is not whether puppies differ. It is whether the differences observed at seven weeks are the same differences that will matter at two years.

2. The Longitudinal Evidence

2.1 The Design

A cohort of Border Collies was followed from the neonate stage to adulthood. Ninety-nine puppies aged two to ten days were assessed for activity, vocalization when isolated, and sucking force. At forty to fifty days, 134 puppies — including ninety-three of those tested as neonates — completed a puppy test at their breeders' homes. All were placed as pets, and fifty participated in a behavioral test at one and a half to two years with their owners (Riemer, Müller, Virányi, Huber & Range, 2014).

This is the design the question requires: the same individuals, measured at three points, across the full developmental span.

2.2 The Result

Linear mixed models found little correspondence between individuals' behavior in the neonate, puppy and adult tests. Exploratory activity was the only behavior significantly correlated between the puppy test and the adult test (Riemer et al., 2014).

One behavior, out of everything measured.

2.3 The Neonate Data

The neonate assessments fared worse still. There was a lack of correspondence between the behavior of neonates and the same dogs in the puppy and adult tests, which the authors describe as implying a lack of validity of that tool for making predictions about future behavior (Riemer et al., 2014).

Some practitioners use neonate tests to complement later assessments when selecting working prospects. The authors note that these have not been validated — and their data indicate why.

2.4 Why This Design Is the Right One

Most of what circulates on puppy testing is cross-sectional or anecdotal — a breeder reporting that their assessments have worked out well, or a study correlating puppy scores with something measured shortly afterwards.

Following the same individuals across the developmental span is the only design that answers the question, because the claim is about individual trajectories rather than group averages (the same requirement that other prediction questions face).

2.5 What the Study Cannot Show

Fifty adults from an initial 134 is attrition, and the dogs were a single breed from a specific population. Whether a different test battery, applied at a different age, would perform better is not addressed.

What the study establishes is that this design — the one that answers the question properly — found almost nothing carrying over.

3. Where the Tests Came From

3.1 The Origin

Many commonly used puppy test batteries trace back to protocols developed from the 1970s onward, at a time when validation against later outcomes was not a routine step in developing behavioral instruments. Their structure draws on the developmental framework established at the Jackson Laboratory (Scott & Fuller, 1965) and on clinical observation.

They were proposals for how to assess a puppy, and they entered practice before anyone established what the assessment predicted.

3.2 Why That Order Matters

A test that enters practice before validation acquires a constituency. Breeders build placement procedures around it, programs build selection around it, and by the time the evidence arrives it is competing against an installed practice rather than against an open question.

The pattern is the same one that produced other unvalidated constructs in this field (with the mechanism documented separately).

3.3 The Vocabulary Problem

Several of the original scales use terms — dominance, submissiveness, social attraction — carrying theoretical commitments that have not survived. An elevation test scored as "dominance" imports a framework the behavioral literature has since abandoned (as the dominance evidence sets out).

That does not make the observation worthless. It does mean the score name is a hypothesis rather than a description.

3.4 What Came After What

Common puppy temperament protocols were widely adopted before their predictive validity had been established. The studies in the following sections were conducted afterwards, testing a practice that already existed — which is the reverse of how a diagnostic instrument is normally developed.

4. The Meta-Analysis

4.1 What It Pooled

Thirty-one studies reporting temporal consistency of dog personality were identified and combined. The overall estimate suggested substantial consistency at r = 0.43 across all ages and intervals (Fratkin, Sinn, Patall & Gosling, 2013).

That figure is worth stating first, because the finding is not that dogs are inconsistent. It is that consistency depends heavily on when you measure.

4.2 The Age Effect

Estimates were categorized by whether dogs were under or over twelve months at first testing. Both were significantly different from zero and significantly different from each other.

The average weighted adult consistency estimate was r = 0.51, against r = 0.30 for puppies — 1.7 times as large (Fratkin et al., 2013).

4.3 Which Traits Hold and Which Do Not

This is the result that matters most for practice. In puppies, aggression and submissiveness were the most consistent dimensions, while responsiveness to training, fearfulness and sociability were the least consistent (Fratkin et al., 2013).

In adult dogs there were no dimension-based differences in consistency at all.

4.4 Why That Ordering Is Awkward

Read the second list again. Trainability, fearfulness and sociability are exactly what a family wants predicted and what a working program screens for. They are the least stable dimensions at the age tests are administered.

The dimensions that do hold — aggression and submissiveness — are the ones a seven-week-old has least opportunity to display in a standard battery (and aggression is situation-specific in any case).

4.5 What Improved Consistency

Three moderators. Consistency was higher when the interval between assessments was shorter, when the measurement tool was exactly the same on both occasions, and when dogs were older at first measurement (Fratkin et al., 2013).

All three run against puppy testing, which uses a long interval, a different instrument at follow-up, and the youngest possible first measurement.

4.6 What r = 0.30 Means

A correlation of 0.30 is not zero, and it is worth being precise about what it does and does not permit.

At the group level it is a real association: across many puppies, higher scores on a dimension go with somewhat higher scores later. At the individual level it means the score accounts for roughly nine percent of the variance in the later measurement — which leaves the great majority of what a particular dog will be unexplained.

Group-level association and individual prediction are different claims, and the gap between them is where puppy testing operates (the same distinction that governs breed-based inference).

4.7 The Finding Nobody Quotes

There was no difference in consistency between dogs tested first as puppies and later as adults, and dogs tested first as puppies and later again as puppies (Fratkin et al., 2013).

That is unexpected and informative. The problem is not primarily that development intervenes between the two measurements. Something about measuring puppies produces less consistent estimates regardless of when the second measurement occurs.

5. The Service-Dog Data

5.1 The Dataset

At the Swedish Dog Training Centre, 1,235 eight-week-old German Shepherd puppies were tested between 1978 and 1983 across ten score groups. Most were also tested as adults at 450 to 600 days using the established service-dog selection regimen, leaving 867 dogs with complete records from both tests and 630 in the years analyzed for correspondence (Wilsson & Sundgren, 1998).

5.2 The Result

Correspondence of puppy test results to performance at adult age was negligible, and the puppy test was therefore not found useful in predicting adult suitability for service dog work (Wilsson & Sundgren, 1998).

5.3 Why This Dataset Carries Weight

Three features make it unusually informative. The dogs were purpose-bred at one center, which removes much of the environmental variation that would otherwise obscure a signal. The adult test was an established operational selection procedure rather than a research construct. And the sample is large by the standards of this literature.

If a puppy test were going to predict anywhere, a standardized population with an established adult outcome measure is where it would show.

5.4 The Paradox That Makes It Interesting

Heritability was medium-high to high for the behavioral characteristics measured in the puppy test (Wilsson & Sundgren, 1998).

So the test was measuring something real, genetically influenced and reliably scored — and that something was not what determined adult suitability. This rules out the easy explanation. The failure is not that puppy tests measure noise.

5.5 What That Leaves

Two possibilities the data cannot separate. The heritable traits measured at eight weeks may be genuinely different traits from those that matter in an adult working dog. Or they may be the same underlying traits expressed so differently at the two ages that a test designed for one cannot capture the other.

Either way, heritability at eight weeks does not deliver prediction at eighteen months (with the general gap between genetic influence and individual prediction documented elsewhere).

6. Why Prediction Fails

6.1 Development Has Not Finished

A seven-week-old is inside the socialization period, with major developmental changes still ahead — in social behavior, exploration and responsiveness — and most of its formative experience yet to occur (with the developmental sequence set out separately).

Measuring a system mid-construction and expecting the reading to describe the finished structure is optimistic in a specific way: it assumes the remaining construction adds detail rather than changing shape.

6.2 The Environment Contributes Substantially

Between the test and adulthood sit eighteen months of experience — a household, a training approach, a socialization history, illnesses, adolescence. Adult behavior emerges from the interaction between inherited predispositions and that experience, which means early measurement is being asked to predict across a period during which many relevant influences on adult behavior occur.

6.3 A Single Session Is a Single Sample

The situation-specificity problem applies here as it does everywhere in behavioral assessment. One session, in one place, on one day, with one handler, samples a single combination of conditions.

The consistency moderators point the same way: shorter intervals and identical instruments improve estimates (Fratkin et al., 2013), which is a statement about measurement error as much as about development.

6.4 The Traits Are Not Yet Differentiated

A plausible reading of the dimension-level results is that adult personality dimensions have not fully separated at seven weeks. In adults, no dimension was more consistent than another; in puppies, some were and some were not (Fratkin et al., 2013).

That pattern is consistent with a structure still forming rather than with a formed structure measured badly.

6.5 The Test Situation Is Novel for a Puppy

Everything in a standard battery is unfamiliar to a seven-week-old: the room, the tester, the objects, the handling. Which means the score partly measures response to novelty rather than the named trait.

For an adult dog, a novel test situation is one context among many it has encountered. For a puppy, novelty is close to the entire experience — and a measure of how an animal responds to something for the first time is not obviously a measure of a stable disposition (with the arousal complication running alongside).

6.6 Litter Order and Immediate History

Puppies are usually tested in sequence, which means some have been separated from the litter for longer, some have just eaten and some have not, and later-tested puppies may be responding to a tester who has warmed up or grown tired.

These are ordinary sources of measurement variance and they are large relative to the differences the test is trying to detect.

6.7 The Age Gradient Within the Puppy Period

An earlier guide-dog study makes the timing point with unusual precision. Comparing puppy fear-test scores with adult fearfulness, no significant correlations were found for tests conducted at six or seven weeks — while one of three tests at eight weeks and two of four at ten weeks did correlate significantly with adult fearfulness (Goddard & Beilharz, 1984).

Adult fearfulness could be predicted to some degree from fearfulness at three months, and the accuracy of prediction improved with age (Goddard & Beilharz, 1984).

That is the same conclusion the meta-analysis reaches from pooled data, visible here as a gradient within a few weeks: prediction is not absent and it is a function of when the measurement is taken.

6.8 Adolescence Sits in Between

Whatever is measured at seven weeks has to survive puberty to appear in an adult test. Responsiveness to the caregiver measurably dips and recovers during that period (as the adolescence evidence documents), which is one more source of change between the two measurements.

7. What the Tests Do Measure

7.1 The Current State

A puppy test is a valid description of how this puppy behaved in this situation today. That is not nothing, and it is not what the test is used for.

7.2 Exploratory Activity

The one behavior that carried over from puppy to adult test was exploratory activity (Riemer et al., 2014), which is also among the simplest things measured — how much the animal moves and investigates in a novel space.

It is worth noticing which behavior survived. Not sociability, not fearfulness, not trainability, but general activity in a new environment.

7.3 Why the Surviving Behavior Is Telling

Exploratory activity is the least interpreted thing a battery measures. It requires no inference about motivation, no judgment about what a response means, and no theoretical commitment in the scoring.

The behaviors that failed to carry over are the ones requiring interpretation — whether a puppy's response to restraint indicates dominance, whether hesitation indicates fearfulness. That pattern suggests part of the problem sits in the inference rather than in the animal (with the same issue across behavioral constructs).

7.4 Aggression and Submissiveness

The meta-analysis found these the most consistent dimensions in puppies (Fratkin et al., 2013). Whether a standard puppy battery can elicit them meaningfully in a seven-week-old is a separate question, and the answer is probably not.

7.5 What Owners Actually Observe

Owners of adult dogs frequently report that the puppy they chose "was always like this", and the recollection is not evidence.

A description supplied at eight weeks becomes the frame through which eighteen months of behavior is interpreted, and confirming instances are more memorable than disconfirming ones. That is ordinary, and it is why owner recollection cannot substitute for the longitudinal data (with the same limitation across owner-reported measures).

7.6 Immediate Practical Information

There is a legitimate use that requires no prediction at all. A puppy that reacts strongly to a noise today needs that handled during placement this week, whether or not it says anything about the dog at two years.

Assessment for present management is defensible. Assessment for prognosis is what the evidence does not support (a distinction that applies to behavioral testing generally).

8. Reliability, Validity and Prediction

8.1 Three Questions Routinely Merged

The same separation that clarifies shelter assessment applies here, and it explains why a test can be well-constructed and still useless for its stated purpose.

Reliability asks whether two assessors scoring the same puppy agree, and whether the same puppy scores similarly a week later. Construct validity asks whether the battery measures the trait it names. Predictive validity asks whether today's score forecasts adult behavior.

8.2 Where Puppy Tests Actually Stand

Reliability is generally adequate — the batteries are structured and the scoring is defined. Construct validity is uncertain, since a seven-week-old's response to elevation may or may not be an index of anything called dominance.

Predictive validity is where the evidence is clearest, and it is poor.

8.3 Why the First Two Cannot Rescue the Third

A test can be perfectly repeatable and measure exactly what it claims, and still fail to predict — because prediction requires the measured property to remain stable over the interval concerned.

That is the crux. Puppy tests do not fail because they are badly built. They fail because the measured characteristics are not sufficiently stable over the developmental period relevant to adult outcomes.

8.4 What Would Change the Answer

Longer or repeated assessment, closer to the outcome, using the same instrument. Every one of those moves the test away from what a puppy test is — a single early session — and toward what the consistency data actually favor.

9. What Selection Can Use

9.1 Heritability Is Not the Problem

The service-dog data show behavioral traits with medium-high to high heritability at eight weeks (Wilsson & Sundgren, 1998), while the Swedish program relied on adult index values rather than puppy scores for its selection decisions.

Genetic selection works. It requires traits that are measured in relation to the outcome being selected for — which, for adult working suitability, means measurement at an age where the two are related.

9.2 Parents and Relatives

Where prediction from the individual puppy fails, prediction from the pedigree does not. Heritable traits assessed in adult relatives carry information that a seven-week test does not, and it is available before the litter is born.

9.3 Later Assessment

Working programs that succeed do so by testing at five to eight months and beyond, not at seven weeks — with predictive accuracy improving substantially over that range.

That is the same conclusion the consistency data reach from the other direction: older is better, and the improvement is large.

9.4 Why Later Testing Costs Less Than It Appears

The objection to assessing at six months is that the program has already invested in raising the dog. That calculation looks different when the alternative is selecting on a measure that does not predict.

A puppy test that identifies the wrong candidates costs the full training investment in dogs that will not qualify, plus the loss of dogs wrongly excluded. Testing later costs more per assessment and less per successful placement.

9.5 Rearing Conditions

The variable a breeder controls most directly is the environment before placement, not the selection of which puppy goes where. Early experience shapes what a test would later measure (with the developmental basis set out separately).

10. Why the Practice Persists

10.1 It Delivers What People Want

A profile at seven weeks is exactly what a family, a breeder and a program all want at the moment they want it. The alternative — wait, observe, decide later — is unavailable when the litter is being placed.

Demand for a prediction does not create the ability to make one, and it reliably creates the practice (a pattern documented for other confident constructs).

10.2 It Confirms Itself

The dog placed as the confident one is raised by people who were told it was confident, and it is often confident. Expectation shapes handling, handling shapes behavior, and the outcome reads as vindication.

Nothing in a normal placement distinguishes prediction from self-fulfilment.

10.3 Failures Are Attributed Elsewhere

A test that predicted well is a good test. A test that predicted badly is a dog that was raised wrong, or a litter that was tested on a bad day. The framework survives either outcome.

10.4 The Results Feel Substantive

Numerical scores and named categories carry the appearance of measurement, and the batteries are structured enough to be reliable — which is exactly the property that makes them convincing without being informative (the reliability–prediction gap again).

10.5 The Prediction Is Never Checked

The structural reason the practice survives is that almost nobody follows up. A breeder places eight puppies and hears from perhaps half of them, at intervals, in unstructured form.

Verifying a prediction requires systematically comparing forecast against outcome across a cohort. That is a research design, and it is not what happens when a litter goes to homes.

10.6 The Real Function May Be Different

There is a defensible role that is not prediction. Running a litter through structured exposures gives puppies novel experiences during the socialization period, gives the breeder a systematic look at each animal, and gives new owners a shared vocabulary.

Those are reasonable purposes. They do not require the results to forecast anything.

11. What This Means in Practice

11.1 For Breeders

Test if it is useful for structuring observation and exposure. Do not report the results as predictions, and do not place on the strength of them.

Pedigree information on adult relatives carries more signal than the seven-week battery, and rearing conditions carry more still.

11.2 For Buyers

A puppy described as "the confident one" or "the shy one" has been described accurately as of that morning. Sociability, fearfulness and trainability were the least consistent dimensions in the meta-analysis (Fratkin et al., 2013), which is precisely what those labels refer to.

Ask instead about the parents' adult behavior, the rearing environment, and what the puppies have been exposed to.

11.3 For Working Programs

The Swedish data are unambiguous: puppy testing did not predict service-dog suitability, and selecting on adult index values did improve the population (Wilsson & Sundgren, 1998).

Testing later costs more per dog and produces information that exists.

11.4 For Behavior Professionals

Two things follow for anyone advising on placement or assessing a young dog.

A puppy test result in a dog's history is a record of one morning, not a baseline. Treating it as a starting point against which later behavior is compared imports a prediction the data do not support.

And where owners arrive with a label — "we were told he was the dominant one" — the label is worth unpicking rather than working around. It shaped eighteen months of handling before the consultation (with expectation effects documented across owner-reported measures).

11.5 What Not to Conclude

That puppies are blank slates. Behavioral traits at eight weeks were substantially heritable (Wilsson & Sundgren, 1998), individual differences are real, and the pedigree matters.

The finding is narrower and more specific: the differences observable in an individual puppy on one occasion do not forecast that individual's adult behavior, which is a claim about measurement rather than about the animal.

12. What Would a Better Instrument Look Like

12.1 The Design Implied by the Data

The consistency findings are not only a verdict. They specify what would work better, because each moderator points at a design feature.

Consistency rose with shorter intervals, identical instruments and older age at first measurement (Fratkin et al., 2013). An instrument built on all three would assess later, repeat the same battery at follow-up, and shorten the gap.

12.2 Repeated Rather Than Single

The clearest implication is that one session is the wrong unit. Averaging across occasions reduces measurement error, and the same logic has been demonstrated elsewhere in canine assessment — where averaging a measure across recordings predicted an outcome better than any single reading.

A breeder who observes across eight weeks has better information than a battery run once, and rarely records it systematically.

12.3 Target the Dimensions That Hold

Aggression and submissiveness were the most consistent puppy dimensions (Fratkin et al., 2013). Whether a battery can elicit them meaningfully at seven weeks is doubtful, and it is the direction a validation study would take.

12.4 Predict Something Specific

Most existing work asks whether puppy tests predict personality dimensions or working suitability. Neither is what a pet owner needs.

Whether early assessment predicts specific outcomes — noise sensitivity, separation-related behavior, resource guarding — is barely addressed, and those are narrower questions with clearer criteria (each of which has its own literature).

12.5 Why Nobody Has Built It

Such an instrument would be expensive, slow, and would deliver its answer after the age at which puppies are placed. The practical value of a puppy test lies in its timing, and the timing is the problem.

That is not a failure of research effort. It is a structural conflict between when the information is wanted and when it becomes available.

13. Predictive Validity Depends on the Outcome

13.1 The Question Is Incomplete Without It

"Do puppy tests predict?" cannot be answered as stated, because prediction is always prediction of something. A test can be informative for one outcome and useless for another, and the literature has concentrated on a narrow set.

13.2 What Early Assessment Plausibly Describes

Current handling tolerance — whether this puppy can be picked up, restrained and examined today. Directly observable, immediately useful, no prediction required.

Momentary reactivity — how this puppy responded to a noise or a novel object this morning.

Within-litter differences — which puppy was more active than its siblings, in this session. That is a relative judgment within a controlled comparison, and it is the most defensible thing a battery produces.

13.3 What It Does Not Predict

Adult aggression, trainability, human-directed sociability and working suitability — the four outcomes that placement and selection decisions actually turn on (Riemer et al., 2014; Fratkin et al., 2013; Wilsson & Sundgren, 1998).

13.4 Why the Distinction Matters

A breeder who says "this puppy handles restraint well today" is making a claim the test supports. One who says "this puppy will be easy to handle" is making a claim it does not.

The two sound similar and differ entirely in what would falsify them.

13.5 Personality Exists — the Measurement Is the Problem

None of this argues that dogs lack stable individual differences. Personality in dogs is well established, and a shyness–boldness dimension assessed at twelve to eighteen months predicted performance in working trials across 2,655 German Shepherds and Belgian Tervurens, with bolder dogs also reaching success at a younger age (Svartberg, 2002).

That is prediction from a personality measure, and it worked — at twelve to eighteen months, not at seven weeks. The constraint identified throughout this article is about when the measurement is taken, not about whether there is anything there to measure (with the personality construct set out separately).

13.6 The Methodological Backdrop

A review of temperament and personality research in dogs concluded that the field's instruments varied widely in what they measured and how well they had been evaluated (Jones & Gosling, 2005), which is the broader context in which puppy batteries sit.

14. Summary at a Glance

One behavior carried over in the longitudinal study — Exploratory activity was the only behavior significantly correlated between the puppy test and the adult test in a Border Collie cohort followed from the neonate stage (Riemer et al., 2014).

Neonate tests fared worse — Little correspondence between neonate behavior and the same dogs later, implying a lack of validity for prediction (Riemer et al., 2014).

Puppy consistency is about half the adult level — r = 0.30 against r = 0.51 across 31 studies (Fratkin et al., 2013).

The least consistent puppy traits are the ones people test for — Responsiveness to training, fearfulness and sociability; aggression and submissiveness were the most consistent (Fratkin et al., 2013).

Development is not the whole explanation — Consistency was no different for puppy-to-adult than for puppy-to-puppy comparisons (Fratkin et al., 2013).

One of the largest service-dog datasets found correspondence negligible — Across 630 German Shepherds tested at eight weeks and again at 450–600 days (Wilsson & Sundgren, 1998).

The traits were nonetheless heritable — Medium-high to high heritability for the puppy test characteristics, which rules out the explanation that the test measured noise (Wilsson & Sundgren, 1998).

Predictive validity was established after adoption, not before — The batteries entered practice and the evidence followed.

Prediction improves with age even within the puppy period — No significant correlations with adult fearfulness for tests at six or seven weeks; some at eight and ten weeks; better still at three months (Goddard & Beilharz, 1984).

Personality measured later does predict — A shyness–boldness dimension assessed at 12 to 18 months predicted working-trial performance across 2,655 dogs, with bolder dogs succeeding younger (Svartberg, 2002).

Selection ran on adult phenotypes — The same program used adult index values rather than puppy scores for selection decisions (Wilsson & Sundgren, 1998).

15. Research Gaps and Critical Appraisal

The longitudinal study is one breed and fifty adults. Riemer et al. (2014) is the right design and a modest sample, in Border Collies, with attrition from 134 puppies to 50 adults.

The meta-analysis pools heterogeneous instruments. Thirty-one studies using different batteries, intervals and outcome measures. That is what a meta-analysis does, and it means the pooled figure describes a literature rather than a test.

The service-dog data are historical and breed-specific. German Shepherds bred at one center between 1978 and 1983, tested with a protocol of that period.

No study has tested an optimised puppy battery. Whether a test designed with the consistency findings in mind — targeting the dimensions that do hold, at the oldest feasible age — would perform better is unknown, because nobody has built and validated one.

Prediction has mostly been assessed against the wrong outcomes. Working suitability and personality dimensions are what has been measured. Whether early tests predict specific problem behaviors, which is what most owners care about, is barely addressed.

Placement matching has barely been studied. Nearly all the evidence concerns working suitability or personality dimensions. Whether puppy tests improve owner satisfaction or reduce relinquishment — the outcome that matters for pet placement — has not been examined.

Self-fulfilment has not been controlled for. No study has blinded owners to their puppy's test results and compared outcomes, which is the design that would separate prediction from expectation.

The alternative has not been quantified either. That pedigree and rearing carry more information than the puppy test follows from the heritability data and general developmental evidence, and the comparison has not been made directly.

16. Conclusion

The evidence on puppy temperament testing is unusually consistent for this field, and it points one way. A cohort of Border Collies followed from the first days of life to adulthood produced a single behavior that carried over from the puppy test to the adult one, and that behavior was general exploratory activity. A meta-analysis of thirty-one studies found personality consistency in puppies at roughly six-tenths of the adult level, with the lowest estimates for responsiveness to training, fearfulness and sociability — which is a list of the three things a puppy test is used to assess. And one of the largest working-dog datasets available found correspondence between eight-week scores and adult service suitability negligible across six hundred and thirty German Shepherds. That last study also rules out the comfortable explanation. The traits measured at eight weeks were substantially heritable, so the test was picking up something real and genetically influenced; it simply was not the thing that determined what the adult dog could do. What follows is narrower than "puppy tests are worthless" and more useful. The tests describe a puppy accurately on the day, which supports placement management and structured early exposure. What they do not support is prognosis — and the alternatives are available: adult behavior in the parents, the rearing environment, and assessment at an age where the consistency data say the measurement starts to hold (much as other single-occasion assessments carry less information than the decisions made on them assume).

Key Insights

  • The longitudinal study found one behavior that carried over. Border Collies assessed as neonates, as puppies at 40–50 days, and as adults at 1.5 to 2 years showed little correspondence across the three tests, with exploratory activity the only behavior significantly correlated between the puppy and adult assessments (Riemer et al., 2014). Neonate tests showed no correspondence at all.
  • Puppy consistency is roughly half the adult level, and lowest where it matters. Across 31 studies, consistency was r = 0.30 in puppies against r = 0.51 in adults — and within puppies, responsiveness to training, fearfulness and sociability were the least consistent dimensions (Fratkin et al., 2013). Those three are what families and programs want predicted.
  • Development is not the whole story. Consistency was no different when dogs were tested as puppies and again as adults than when they were tested as puppies twice (Fratkin et al., 2013). Something about measuring puppies produces less consistent estimates regardless of the interval, which points at measurement as well as at change.
  • The heritability result rules out the easy explanation. Across 630 German Shepherds, puppy test traits showed medium-high to high heritability while correspondence with adult service-dog performance was negligible (Wilsson & Sundgren, 1998). The test was measuring something real and genetically influenced — just not the thing that mattered later.
  • A correlation of 0.30 permits group description, not individual prediction. It accounts for roughly nine percent of the variance in the later measurement, which leaves most of what a particular puppy will become unexplained — and individual placement decisions are exactly what puppy tests are used for.
  • The improvement with age is visible within weeks, not just years. Guide-dog fear tests at six and seven weeks showed no significant correlation with adult fearfulness, while some tests at eight and ten weeks did, and prediction from three months was better again (Goddard & Beilharz, 1984). The gradient is the finding: this is a question of timing rather than of whether puppies differ.
  • Personality exists and can be measured — later. A shyness–boldness dimension assessed at 12 to 18 months predicted performance in working trials across 2,655 German Shepherds and Belgian Tervurens (Svartberg, 2002). The problem identified here is the age of measurement, not the existence of stable individual differences.
  • The alternatives exist and are not exotic. The same Swedish program relied on adult index values rather than puppy scores for its selection decisions (Wilsson & Sundgren, 1998). Adult behavior in relatives, the rearing environment, and assessment at an older age all carry information that a seven-week battery does not.

References

Fratkin, J. L., Sinn, D. L., Patall, E. A., & Gosling, S. D. (2013). Personality consistency in dogs: A meta-analysis. PLoS ONE, 8(1), e54907. https://doi.org/10.1371/journal.pone.0054907

Goddard, M. E., & Beilharz, R. G. (1984). A factor analysis of fearfulness in potential guide dogs. Applied Animal Behaviour Science, 12(3), 253–265. https://doi.org/10.1016/0168-1591(84)90118-7

Jones, A. C., & Gosling, S. D. (2005). Temperament and personality in dogs (Canis familiaris): A review and evaluation of past research. Applied Animal Behaviour Science, 95(1–2), 1–53. https://doi.org/10.1016/j.applanim.2005.04.008

Riemer, S., Müller, C., Virányi, Z., Huber, L., & Range, F. (2014). The predictive value of early behavioural assessments in pet dogs — a longitudinal study from neonates to adults. PLoS ONE, 9(7), e101237. https://doi.org/10.1371/journal.pone.0101237

Scott, J. P., & Fuller, J. L. (1965). Genetics and the Social Behavior of the Dog. University of Chicago Press.

Svartberg, K. (2002). Shyness–boldness predicts performance in working dogs. Applied Animal Behaviour Science, 79(2), 157–174. https://doi.org/10.1016/S0168-1591(02)00120-X

Wilsson, E., & Sundgren, P.-E. (1998). Behaviour test for eight-week old puppies — heritabilities of tested behaviour traits and its correspondence to later behaviour. Applied Animal Behaviour Science, 58(1–2), 151–162. https://doi.org/10.1016/S0168-1591(97)00093-2