top of page

Why I Use Kaufman (K-SLP) (Alongside DTTC), Even Though DTTC Has a More Established Evidence Base

Aug 28
15 min read

If you've heard much of my teaching about childhood apraxia of speech (CAS), you've probably noticed that I talk a lot about syllable shapes.

Probably more than most people.

And that's because I think one of the most important questions we can ask as clinicians is this:

How do humans actually learn to speak?

Before we start debating DTTC versus Kaufman, blocked versus distributed practice, or which intervention has the strongest evidence base, I think it's worth stepping back and asking a much more fundamental question:


How does speech develop in the first place?

How do children learn speech?



One explanation that profoundly influenced my thinking came from David Hammer.

He encouraged me to think about the first words babies typically say.

What are they?

Mama.

Papa.

Baba.

Why these words?


It's not because babies have carefully analysed which words are most functional.


In fact, many highly functional words are much harder to say. Imagine expecting a typical 12-month-old to begin with words such as chicken, thank you, yellow or cheese.

Instead, babies begin with speech patterns that are accessible to their developing speech system.


The sounds /m/, /b/ and /p/ are early-developing sounds made with the lips.

And the vowel /a/ requires very little precision. At its simplest level, it can be produced with a relatively rudimentary jaw-opening movement.


Put those together and you get:


mama

papa

baba


Those early words aren't random.


They reflect some of the earliest speech templates available to a developing speech system.

This aligns beautifully with Marilyn Vihman's work describing how children develop speech through templates and syllable shapes.


Children don't learn speech one word at a time.

They learn patterns.

Templates.

Syllable shapes.

And once those templates become available, they begin using them across many different words.

The shape comes first.

The refinement comes later.


And the more I have thought about childhood apraxia of speech, the more convinced I have become that this principle should sit at the centre of our clinical thinking.



Speech Doesn't Stop Being Pattern-Based Once We're Adults


Something that fascinates me is that the importance of syllables and speech patterns doesn't disappear once speech is acquired.

In fact, psycholinguistic models suggest that highly practised syllables may become stored as planning units.


This idea is often referred to as the Mental Syllabary hypothesis.


The basic idea is that, through years of speaking, we become extremely efficient at planning speech sequences we use frequently.


Rather than constructing every word completely from scratch every time we speak, frequently used syllables become increasingly easy to access and produce.


As Julia Chauvet recently summarised, speakers become faster and more accurate when producing frequent sequences of speech sounds, a phenomenon known as the syllable frequency effect. She also notes that in everyday speech we repeatedly reuse a relatively small number of syllables and that frequently practised syllables become easier to plan. Researchers have proposed that with practice, speakers store motor plans for frequently used syllables within a “mental syllabary”.


This doesn't mean we store every word as a fixed unit.

Nor does it mean speech is simply a collection of memorised syllables.

But it does suggest that speech remains fundamentally pattern-based.

Even in adulthood.


That matters because it supports an idea I've believed for a long time:


Speech isn't organised around individual words. Speech is organised around patterns.

And if speech is fundamentally pattern-based, then helping a child with CAS establish robust speech patterns and syllable shapes becomes incredibly important.



Why this matters for CAS


The more I have thought about childhood apraxia of speech, the more important this idea has become. Because,


I don't think my primary job is teaching words. My primary job is helping a child build a speech system.

And that speech system is built from syllable shapes.


When I'm assessing a child with severe CAS, I'm constantly thinking:


  • Which syllable shapes are available?

  • Which syllable shapes are emerging?

  • Which syllable shapes are missing?

  • Which syllable shapes need strengthening?


That's often much more important to me than whether an individual word is perfect.


The Goal Was Never “Bubble”



Let's take the word:

bubble


Imagine a child says:

buhbah


Many clinicians immediately focus on the errors.

And yes, it isn't the adult production.

But that's often not my first thought.


My first thought is:

Do they have the CVCV template?


Because if they can reliably say:

buhbah


then something important has already happened.


They've established:

✅ two syllables

✅ a CVCV template

✅ a stable word structure

✅ a meaningful production that can be used functionally


And honestly, I would much rather hear a child say:


“buhbah”


than:


“uhbah.”


Why?


Because buhbah contains the CV₁CV₂ template.


The child now has access to the underlying word structure.


Uhbah doesn't.


When we insist on perfect vowels before allowing a child access to the template, we can sometimes leave them stuck producing forms that are missing a critical piece of the structure we're ultimately trying to teach.
And there's another practical problem.Children don't only practise words in therapy. They practise them all day long.

If a child leaves therapy saying:

“uhbah”

then they may produce that same incomplete form dozens or hundreds of times throughout the week.

Every time they talk about bubbles.

Every time they request bubbles.

Every time they use the word at home, preschool or school.

They're repeatedly reinforcing a word form that is missing part of the template.


Personally, I'd much rather they leave therapy saying:

“buhbah.”


Because now they're repeatedly rehearsing a meaningful CVCV shape that I ultimately want them to own.


Once that shape becomes strong and stable, I can start refining it:

buhbah → bubbah → bubble


The vowels can be refined.


The consonants can be refined.


The accuracy can be refined.


But first I want the child to establish the shape itself.


That matters.


A lot.


Because if the child doesn't yet have the shape, there isn't much to refine. But once the shape is stable, now we can begin working on the details. That's why I don't see approximations as the destination. I see them as a bridge.


Why I continue to value Kaufman principles


This is one reason I continue to use aspects of the Kaufman approach.


Not because I believe approximations are the goal.


Quite the opposite.


The adult target is always the goal.


But sometimes the child doesn't yet have access to the adult form.


If a child can't yet produce “bubble”, but can produce “buhbah”, then I now have a CVCV shape that I can strengthen, stabilise and build upon.

And here's where I think Kaufman is sometimes misunderstood.


I've occasionally seen people treat the approximations on the cards as though they are rigid instructions. That was never my understanding of Nancy Kaufman's teaching.


The approximations on the cards are examples.


Suggestions.


A starting point.


The clinician's job is to determine the best approximation for the individual child sitting in front of them.


Then, importantly:

upgrade it as soon as possible.


The approximation exists only for as long as the child needs it.

The moment a more accurate production becomes available, we move forward.


That's exactly what I do clinically.



Not All CVCV Words Are The Same


One of the reasons I continue to value the KSPT framework is that it taught me to analyse syllable shapes in much greater detail.


Not all CVCV words are the same.


Not all CVC words are the same.


And this is one of the biggest errors I see clinicians make.


It's easy to look at a group of words and think:“They're all CVCV.”

But clinically, different consonant transitions, vowel patterns and sequencing demands can create very different learning challenges.


Before the DEMSS existed, The Kaufman Speech PraxisTest encouraged clinicians to think much more specifically about these distinctions.

Different CVCV forms.


Different CVC forms.


Reduplicated patterns.


Variegated patterns.


Front-to-back sequences.


Back-to-front sequences.


Different consonant transitions.


Different vowel transitions.


In my experience, this level of analysis is absolutely critical when working with children with severe CAS.

Children don't simply learn “CVCV.”


They learn specific patterns within those broader categories.


And understanding those patterns often guides my target selection far more than the word itself.


We Are Not Teaching Words


This is probably the most important point in this entire article.


When I use the word:


bubble


I am not primarily interested in teaching the word bubble.


I'm interested in teaching an underlying speech pattern.


Similarly, if I use:


  • paper

  • puppy

  • mummy

  • daddy

  • baby

  • turtle


I'm not really interested in those individual words.


I'm interested in helping a child establish and strengthen underlying speech templates.

The words are examples of the same syllable shape (CV₁CV₂).


The syllable template is the target.

And once we start thinking this way, an interesting question emerges.



Why I'm Interested in Distributed Practice


If the goal is the template rather than the individual word, then we have to ask:


What helps children learn that template most efficiently?

Suppose a child has access to a particular syllable shape.


Should they only practise:

bubble

for weeks?


Or should they also experience multiple words that require related template-level learning?


This is where I become interested in distributed practice.


Not because I think every child should be presented with dozens of targets from day one.

And certainly not because I think blocked practice has no role.


Many children need intensive repetition and support at the beginning.


But once a child begins demonstrating access to a template, I start wondering whether exposure to multiple examples may support broader learning.


After all, if children learn speech through templates, it makes sense that learning may accelerate when they encounter those templates repeatedly across different words.

They're no longer learning bubble.


They're learning a speech template.



A Question I Don't Think We Have Fully Answered Yet


One thing that fascinates me is that DTTC and Kaufman may actually be asking a similar question.


In DTTC, a relatively small number of words are often selected and practised intensively.

But the goal isn't really the words.


The goal is the underlying speech pattern.


Those words are being used as vehicles.


That's a concept I strongly agree with.


Where I become curious is whether there are situations in which broader exposure to a template might support faster generalisation.

If the child has truly learned the template rather than simply memorised a handful of words, then we would expect that pattern to begin appearing in new words over time.


From a motor learning perspective, there are reasons to hypothesise that increased variability and distributed practice may support this process.


But I don't think we currently have the research necessary to confidently answer questions such as:


How many exemplars of a template should we teach?

When should we move beyond a small target set?

Does earlier exposure to multiple examples accelerate generalisation?


Those are questions I would genuinely love to see researched.


Why I Don't Think Current Kaufman Research Answers This Question


The Namasivayam study was an important contribution to the literature and provided promising findings.


But I don't think it answers the questions above.


It was a small Phase I study.


Only five children completed treatment.


The intervention occurred within a boot-camp environment.


And importantly, the directly measured treatment sessions used a two-children-to-one-clinician model.


That makes perfect sense within a research-based intensive programme.


But it isn't the same thing as highly individualised one-to-one therapy.


Anyone who has worked extensively with children with severe CAS knows just how individualised effective treatment can be.


Every cue.


Every prompt.


Every approximation.


Every decision about whether a child is ready for an upgrade.


Every decision about whether to increase complexity.


Every decision about whether the child needs more exemplars or more repetitions.


Those decisions are constantly changing throughout a session.


Personally, I would love to see future research investigating these ideas within highly individualised one-to-one intervention, because I suspect that's where some of the most interesting questions about syllable shapes, target selection and generalisation might be answered.

Taking a Closer Look at the Namasivayam Kaufman Results


Part of what led me to these questions was spending time looking beyond the headline findings in the 2024 Namasivayam study.

It's easy to summarise the results as:

4 of 5 children improved their treated words.

3 of 5 children generalised to untreated words.

But I became much more interested in the individual children behind those numbers.

Participant

Starting Point

Targets

Outcomes

Questions it Raised for Me

Participant 2

3;6, severe CAS, 20% PCC

CV, VC, CVC assimilation

✅ Improved treated words


✅ Improved untreated words


✅ Improved intelligibility


✅ Improved functional communication

Are children learning something bigger than individual words when they learn early syllable structures?

Participant 5

3;1, 9% PCC

CV, VC, CVC assimilation

❌ No significant WWM effect


✅ Improved PCC


✅ Improved intelligibility


✅ Improved functional communication

What should meaningful change look like for a child with such a severely limited speech system?

Participant 4

More advanced targets

/st/ clusters and final /k/

✅ Improved treated words


❌ Did not meet generalisation criterion

Should we expect cluster learning and sound acquisition to generalise in the same way as early syllable-shape learning?

One child, Participant 2, was 3 years 6 months old, had severe CAS and began with only 20% PCC. The targets focused on very early speech structures including CV, VC and CVC assimilation patterns. This child demonstrated significant improvement in both treated and untreated words, alongside broader improvements in intelligibility and functional communication.


That immediately caught my attention.

This wasn't simply a child learning to say a handful of practised words.

There was evidence that learning extended beyond those specific targets.


Another child, Participant 5, was only 3 years 1 month old and began with just 9% PCC.

Nine per cent.

Although this child did not demonstrate a significant treatment effect on the Whole Word Match measure, there were clinically significant improvements in consonant accuracy, intelligibility and functional communication.


I found myself wondering:

What should meaningful change look like for a child starting with such a severely limited speech system?


If a child begins to communicate more effectively, becomes more intelligible and demonstrates improved speech accuracy after only a short group-treatment block, does "non-responder" really capture the whole story?


There was another detail that kept coming back to me as I read the study.


These outcomes were achieved in a dyadic treatment model, with one clinician working with two children with CAS at the same time.


As a clinician, I find that genuinely impressive.


Earlier this year, I attempted to see two children with CAS together and quickly realised just how difficult that is.


CAS therapy requires constant decision-making.


Every few seconds you're evaluating productions, adjusting cues, deciding whether an approximation needs upgrading, determining whether the child is ready for more complexity, monitoring fatigue, judging how many repetitions have occurred and deciding what to target next.


Trying to do that for one child is demanding.


Trying to do it simultaneously for two children with CAS is another level entirely.


That doesn't invalidate the study.


But it does make the outcomes more interesting to me.


When I look at Participant 2 demonstrating significant improvements in treated words, untreated words, intelligibility and functional communication, I can't ignore the fact that these gains occurred while clinician attention was necessarily being shared across two children.

When I look at Participant 5, who began with only 9% PCC, demonstrating meaningful improvements in speech accuracy, intelligibility and functional communication, I can't ignore that those gains occurred within the same treatment model.


Then there was Participant 4.

This child significantly improved the treated words but did not meet the study's generalisation criterion.


When I looked more closely at the actual therapy targets and the untreated probe words, I found myself asking different questions.


The child practised words such as:

stool, stop, stick


but was expected to generalise to words including:

speak, slug and swing.


From a speech-planning and syllable-shape perspective, I don't automatically assume that successful learning of one consonant-cluster sequence should immediately transfer to every other cluster sequence.


And this raises what I think is a really important clinical point.


Not all therapy targets represent the same type of learning.


A child learning early syllable shapes such as CV, VC or CVC assimilation is learning something fundamentally different from a child establishing a specific consonant cluster or a consonant in a particular word position.


Once we move into targets such as /st/ clusters or final /k/, therapy can start to overlap much more with principles traditionally associated with articulation intervention.

When I teach /st/, I don't automatically expect that learning to immediately generalise to /sp/, /sl/ or every other cluster sequence.


Those are different speech patterns.


Likewise, when a child is establishing a new sound such as final /k/, I wouldn't usually expect robust generalisation after learning that sound in only a handful of words.


In my own clinical practice, I would still consider that child to be in the process of establishing the sound before expecting widespread generalisation.


That doesn't make the study's findings invalid.


It simply highlights that different therapy targets may require different expectations regarding acquisition and generalisation.


In contrast, when a child begins generalising early syllable shapes such as CV, VC, CVC assimilation or other well-established sound patterns across untreated words, I think we may be observing a different type of learning altogether.


That's one of the reasons Participant 2 interested me so much.


This child was working on very early syllable structures and demonstrated significant improvement in both treated and untreated words.

To me, that's an interesting demonstration of a child potentially learning something larger than a handful of individual targets.


It brings us back to the question I keep returning to throughout this article:


What exactly is the child learning?


Are they learning words?


Are they learning sounds?


Or are they learning underlying speech patterns and syllable shapes that can begin appearing in entirely new words?


In fact, when I step back and look at the study as a whole, one of the things that stands out to me is that these improvements occurred despite the challenges of a two-children-to-one-clinician treatment model.

Participant 2 demonstrated generalisation.


Participant 5 demonstrated meaningful changes in intelligibility and functional communication.


Participant 4 significantly improved the treated words.


Those gains occurred in a setting where clinician attention and feedback were necessarily shared across two children.


That doesn't allow us to conclude that one-to-one therapy would have produced better outcomes.


We simply don't know.


But it certainly leaves me curious.


Because if meaningful gains can occur within a dyadic treatment model, what might we learn from future studies examining the same principles within highly individualised one-to-one CAS therapy?

That is a question I would genuinely love to see investigated.


Those questions didn't make me doubt the study.


Quite the opposite.


They made me realise that simply counting responders and non-responders may not tell us the whole story.


The more closely I looked at the children, the more I found myself thinking about syllable shapes, target selection, acquisition, practice opportunities and generalisation.


And ultimately, that brought me back to the same question I keep returning to:


What exactly is the child learning?


Because if our real goal is not teaching individual words but teaching speech patterns and syllable shapes, then understanding what drives generalisation becomes one of the most important questions in CAS intervention.


Clinical Experience Is Not The Same As “No Evidence”


There is one final point I think is worth considering.


Clinical experience is not a randomised controlled trial.


It cannot establish treatment efficacy.


It cannot tell us whether one intervention is superior to another.


That is exactly why we need research.


But I also think it is inaccurate to equate decades of accumulated clinical experience with “no evidence.”


The Kaufman approach has been used by CAS clinicians for decades, particularly throughout North America.


The Kaufman Speech Praxis Test (KSPT) was widely used in both clinical practice and research long before the DEMSS existed.


Generations of clinicians have used the framework to analyse speech patterns, select targets and guide intervention.


That history does not prove efficacy.


But nor does it represent nothing.


When experienced clinicians repeatedly observe similar patterns across hundreds or thousands of children, those observations generate hypotheses.

The appropriate scientific response is not:


“Ignore those observations.”


The appropriate response is:


“Let's investigate them properly.”


To me, that is exactly where we currently sit.


The absence of a large, high-quality RCT is not the same thing as evidence that an approach doesn't work.

It means the research still needs to be done.


One example is Jordan LeVan, who has publicly spoken about the role the Kaufman approach played in helping him move from being minimally verbal to verbally communicating.

Of course, an individual story is not a clinical trial and shouldn't be treated as one. But stories like Jordan's remind us that behind every research discussion are real children, real families and real communication journeys.


For me, the most interesting question is:


“What are the active ingredients that help children learn to speak?”
And that's a question I hope future research continues to answer.

The Question I Keep Returning To


The question I keep returning to is:

How do children learn speech? Because,


If speech develops through syllable shapes and templates, then perhaps one of our most important jobs is helping children establish those templates as successfully as possible.

The words matter.


Of course they matter.


But only because they give us a way of teaching something much bigger.


The goal was never bubble.


The goal was helping the child acquire a speech system that eventually supports bubble and thousands of other words.

And that's why I think so much about syllable shapes.



Want to Explore These Ideas Further?


Download my Syllable Shape Hierarchy

A visual overview of the syllable shapes and speech patterns I commonly consider when assessing and treating children with CAS.


Or, if you're interested in diving deeper into CAS assessment, target selection, motor learning and clinical decision-making, you can join the waitlist for the next OKAT CAS Academy cohort commencing February 2027.



Olga Komadina

Speech Pathologist | Olga Komadina Apraxia Therapy




References


Aichert, I., & Ziegler, W. (2004). Syllable Frequency and Syllable Structure in Apraxia of Speech.


Chauvet, J. (2026). How Practice Shapes Motor Planning and Speech Output: Insights From Basic Science. Speech Apraxia International Presentation.


Cholin, J., Levelt, W. J., & Schiller, N. O. (2006). Effects of Syllable Frequency in Speech Production.


Kaufman, N. R. (1995/2005 depending on edition). Kaufman Speech Praxis Test for Children (KSPT). Wayne State University Press (or current publisher information if citing the specific edition you use).


Laganaro, M., & Alario, F. X. (2006). On the Locus of the Syllable Frequency Effect in Speech Production.


Maas, E., Robin, D. A., Austermann Hula, S. N., Freedman, S. E., Wulf, G., Ballard, K. J., & Schmidt, R. A. (2008). Principles of Motor Learning in Treatment of Motor Speech Disorders. American Journal of Speech-Language Pathology, 17(3), 277-298.


Namasivayam, A. K., Cheung, K., Atputhajeyam, B., Petrosov, J., Branham, M., Grover, V., & van Lieshout, P. (2024). Effectiveness of the Kaufman Speech to Language Protocol for Children With Childhood Apraxia of Speech and Comorbidities When Delivered in a Dyadic and Group Format. American Journal of Speech-Language Pathology, 33, 2904-2920.


Schiller, N. O., Meyer, A. S., Baayen, R. H., & Levelt, W. J. M. (1996). A Comparison of Lexeme and Speech Syllables in Dutch.


Vihman, M. M. (2014). Phonological Development: The First Two Years.

 




 
 
 

Comments


bottom of page