A multi-state ABA provider and its software vendor released real-world outcomes data. What the numbers show, and why they want other providers to follow.
Key Takeaways
- Applied behavior analysis has argued about clinical quality for years without organizational-scale outcome data to argue from. The most visible public numbers about the field come from federal audits of improper Medicaid payments, which measure documentation rather than whether children learn.
- Behavioral Framework, whose home-based caseload is roughly 70 percent Medicaid-funded, published six months of skill acquisition data on 1,090 children in partnership with Hi Rasmus, the clinical platform it uses. Both organizations position the paper as a starting point rather than a finished answer, and the provider is already revising the text to clarify how the study was structured.
- The study reports 130,241 skill masteries under its Probe-Teach-Probe workflow against 11,945 under standard programming. Asked about that eleven-to-one gap, the authors said roughly 3.4 of it reflects more service hours delivered and roughly 3.2 reflects a faster rate per hour.
- The Council of Autism Service Providers states that benchmarks for ABA services have not yet been established, and its own benchmarks rest on norm-referenced assessments rather than client-level objectives. Behavioral Framework argues its data speaks to precisely that gap, and is asking other providers to publish comparable numbers.
Brittany Rader kept a spreadsheet for years. Every child who came through Behavioral Framework got a row. “It was a spreadsheet of probably 50 different variables that I was trying to capture for each child,” Rader told Acuity Media Network.
Rader is the President of Clinical Services at Behavioral Framework, which serves Maryland, Virginia, North Carolina, and the District of Columbia. She was chasing two questions at once. The first, whether any individual child was progressing, is what every clinician asks. The second is the one almost nobody in the field can answer with numbers: whether the organization itself is any good.
That second question now has money attached to it. Payers are asking it, state Medicaid programs are asking it, and federal auditors have spent four years producing the closest thing the field has to public data. Rader is blunt about what fills the vacuum. “The field is stuck on talking about quality, nobody knows how to define it, because so many of us are driven by social media posts or what we hear at conferences,” she said.
This month, Behavioral Framework and Hi Rasmus, the Copenhagen-based clinical data platform it runs on, published a 25-page white paper reporting six months of skill acquisition outcomes across 1,090 home-based clients. The numbers reward a close reading, and the authors are the first to say the work has limits. The more durable thing about it may be that it exists at all.
Why ABA Providers Rarely Publish Outcomes Data
The public record on ABA quality is dominated by enforcement. Since 2022, the Department of Health and Human Services Office of Inspector General has conducted a multi-state review of Medicaid ABA payments, identifying at least $56 million in improper payments in Indiana, $18.5 million in Wisconsin, $45.6 million in Maine, and roughly $78 million in Colorado. Those audits, as Acuity has reported, largely turn on paperwork: missing session notes, missing signatures, staff without required credentials. They say a great deal about billing and nothing about learning.
Rader made the same point from the provider side. “The data and the narrative are currently being pulled from the OIG audits,” she said. What providers publish instead tends to be small. “You can create really nice, neat, clean, homogeneous groups out of 25 kids,” she said, describing the typical provider white paper. She does not think those groups resemble the full picture of the work. “99 percent of us are not in sterile clinic rooms.”
Dr. Cate Davis, Director of Research and Outcomes at Hi Rasmus and a co-author, framed the gap in academic terms. “Most published ABA research right now happens in controlled settings, very small settings, single clinics, single cohorts, tightly defined populations,” she told Acuity Media Network. The problem is well documented on the measurement side too: Acuity has covered the field’s unresolved argument over what an ABA outcome even is, and executives at this year’s BHASe Summit conceded that outcomes measurement remains improvised even at sophisticated operators.
The framework at the center of the study did not begin as a research question. It began as a staff training and retention problem. As the Behavior Analyst team grew, each clinician arriving with programming preferences of their own, the company found it could not compare anything to anything. Rader recalled the moment Angela West, the Chief Clinical Officer and company founder, recognized the opportunity: “I need everybody to get on the Probe Teach Probe framework, and this is how we need to do it.” West’s push for standardization carried a dual goal: better clinical outcomes for children, and a workflow her growing team of Board Certified Behavior Analysts could implement consistently without added burden. The rollout ran on paper for a couple of years before the company looked for software that could support it.
Probe-Teach-Probe vs. Standard Programming: What the Skill Acquisition Data Show
Probe-Teach-Probe structures a goal into three phases. Each active target in a goal is presented once without prompts or reinforcement, a procedure the literature calls a cold probe. Targets the learner already knows are set aside. Teaching concentrates on what remains, and a second probe follows. Standard programming, the comparison condition, does not formally separate probe and teaching phases. Targets may be run in various sequences, with or without prompts, leaving the structure of each session largely to clinician discretion.
The headline result is the gap in raw counts. Over six months, PTP programs produced 130,241 skill masteries, compared with 11,945 for standard programming. The paper attributes that difference, more than 118,000 additional masteries, “directly to the instructional efficiency of the PTP framework as operationalized through Hi Rasmus.”
The paper does not report how many service hours went into each condition, but the figure can be derived from its own tables, since program-level rates are masteries divided by hours. The arithmetic yields roughly 102,800 activity hours for PTP, compared with roughly 30,500 for standard programming.
Upon being presented with that calculation, Davis confirmed it and volunteered the correction herself. “While they spent more time in Probe Teach Probe, we would naturally want to see more skills mastered,” she said. “But when you break it down, and you get down to it, it actually is a 3.2x.” In other words, the eleven-to-one figure is roughly 3.4 times the exposure multiplied by roughly 3.2 times the rate, and the second number is the one that describes efficiency.
The rate difference itself is consistent. PTP outperformed standard programming across all six months, achieving monthly rates between 1.171 and 1.366 mastered targets per activity hour, compared to 0.200 to 0.754. The paper reports this as statistically significant, Z = -2.201, p = .028, with a large effect size. That test compares six paired monthly observations, which makes the result a precise statement that PTP led in every month and a more limited one about magnitude. The paper lists this first among its limitations.
A second question concerns what each condition can detect. Because PTP opens with an unprompted probe of every active target and standard programming does not, the two differ in how mastery surfaces, not only in how instruction is delivered. The paper says as much, noting that standard programming “may incorporate prompting immediately following the instructional cue, which can reduce opportunities to assess” independent skill acquisition. Rader described the probe in similar terms. “We knew when we were looking at these probes, these initial probes, it was clean data,” she said.
Rader was candid that some underlying details are off limits, as they were described as the organization’s intellectual property. The paper says goals “best suited to PTP” were assigned to that workflow, which is the kind of selection an archival design cannot correct for, and it says so.
ABA Quality Benchmarks: Why No Standard for Skill Acquisition Rate Exists
None of this is what makes the paper unusual. Buried in its Future Directions section is a sentence describing the field’s actual predicament: there are no empirically established benchmarks for rate of skill acquisition in ABA service delivery. That is not a quarrel with anyone. The Council of Autism Service Providers says the same thing in its own practice guidelines, which state that behavioral health care has not yet established benchmarks for ABA services.
The benchmarks CASP does publish, in a 2025 companion to those guidelines, rest on norm-referenced instruments: expected mean change on the Vineland adaptive behavior composite, on IQ, and on the Childhood Autism Rating Scale, sorted by weekly treatment intensity. Those measure something real, and something different. Rader argues that a provider can look fine against population-level assessment scores while masking meaningful variation in how efficiently instruction is delivered and how quickly a child is actually learning. CASP itself flags that space, calling for short-term outcome indicators shown to predict long-term ones, on the model of blood glucose readings predicting diabetic complications.
Davis, who sits in the CASP outcomes special interest group, said the alignment is the point. “We’re talking about trial-by-trial data being a short-term outcome indicator for maybe those long-term things that we just couldn’t get to in the white paper,” she said. “I think we align very closely with how CASP talks about outcomes.” Similar work exists at scale, notably Linstead and colleagues’ 2017 analysis of objectives mastered by 1,468 children, but it indexes mastery by intensity and duration rather than establishing a rate.
Neither author claims the study settles the question. “This isn’t an experimental design. No controlling or manipulating variables to prove cause and effect. This is a retrospective review of archival data,” Davis said. “We come in with limitations. We come in with this being our starting point.”
Davis pointed to work already underway in that direction, including Dr. Ivy Chong’s argument that standardization must come first because comparisons are meaningless until providers agree on what is being counted. “That strongly aligns with the mission of this work,” she said. The question underneath it, she added, is a practical one: “Can we improve what a session hour looks like, and therefore work our way to value-based care?”
It’s a practical outlook. “Payers and policymakers are coming to organizations and saying, what is the quality of service that you are providing,” Davis said. She noted that six months, the study window, is a common authorization period in payer contracts. As the field edges toward value-based arrangements, a provider able to produce a defensible number has an argument that a provider citing hours delivered does not.
A data platform stands to gain from a field that measures itself, and the paper discloses the arrangement, noting that Behavioral Framework is a Hi Rasmus subscriber and that no author received compensation from the other organization. Davis argues the ledger runs elsewhere. “Who actually benefits from data visibility is the child or the client that we’re working with,” she said. “Having eyes into data impacts every single person who is inherently providing services for a client.” That holds, she said, regardless of the vendor: “I don’t want people to only pull out of this that they have to use Hi Rasmus to have quality outcomes research being pushed out there. Set aside the software.”
Whether other providers follow is the open question. The largest ABA platforms have built proprietary outcomes systems, as Acuity’s survey of the sector’s ten biggest operators documented, and none has published organizational results at this scale. The operational case for doing so is not obscure: maintaining clinical consistency across sites is among the hardest problems in multi-site behavioral health, and it is difficult to manage what nobody counts.
The invitation is the news. “We are not coming to the table saying, look at us, we’ve got the solution,” Rader said. “What we’re doing is we’re trying to fill a gap, and I think if we have other providers’ data at this level, that’s what helps bridge these benchmarks that are missing.” She put the ask plainly: “We are a big provider. We are willing to put this data out there. We think it’s strong data. We encourage others to share their data if they have it, so that we can actually start having tangible information to have this conversation.”
Davis framed it as an opening rather than a verdict. “We’re coming to the table and saying, who else wants to join us in this conversation?” she said. “Who else wants to dig into the data? Who else wants to understand limitations?” A field that has argued about quality in the abstract for a decade now has one provider’s numbers on the table, limitations and all. The second set, from someone else, is what would turn a white paper into a benchmark.






