Seven autism therapy providers independently benchmarked their own outcomes against CASP’s intensity standards. Six shared the results with Acuity, in a test of what transparency can do.
Key Takeaways
- Seven providers independently analyzed their outcomes by the same yardstick. The six that went on the record with Acuity are 360 Behavioral Health, Autism Learning Partners, LEARN Behavioral, Caravel Autism Health, Hopebridge, and Helping Hands Family. Broadly, the group reports, their real-world results tracked the published dose-response.
- The hard part was the data, not the therapy. Dennis Dixon spent forty to sixty hours matching billing records to test scores by hand, because the software was built to care for one child at a time, not to measure a thousand. The effort doubled as a candid look at how much groundwork the field still owes value-based care.
- The finding is real, and the group is clear-eyed about its limits. CASP’s benchmarks describe groups, not individual children, and the clinicians say so plainly. They leaned on one measure, the Vineland-3, that they treat as a shared starting point rather than the last word.
- Underneath sits the question they are trying to answer: what a good outcome is, and how to protect it. With payers weighing hours and audits multiplying, the providers want the field, not the spreadsheet, to define quality. They are alert to the risk that a benchmark meant as a floor becomes a ceiling.
For several weeks earlier this year, one of the most accomplished clinical scientists in autism care spent his evenings doing something closer to archaeology than research.
Dr. Dennis Dixon, PhD, the Chief Clinical Officer of 360 Behavioral Health and an author on some of the most-cited work on how much therapy children with autism actually need, sat between two databases that would not speak to each other. One, a practice-management system called Lumary, knew how many hours each child had received. The other, Pearson’s Q-global, held the test scores meant to show whether those hours had made a difference in the child’s life. The information existed. It had simply never been built to be read together. For the records the software could not reconcile on its own, Dixon matched them by hand, one child at a time, a process that ran to somewhere between forty and sixty hours. “These initiatives don’t just require good clinical practices,” he said, “but also good data practices.”
He was not working alone. Dixon was the most visible hand in an uncommon act of cooperation: seven autism-therapy organizations that ordinarily compete for the same families and clinicians agreed to run the same analysis on their own clients and lay the results side by side. Each measured its own outcomes against a single external benchmark, then chose to make the findings public through Acuity rather than a wire service. The shared motive, in their telling, was less about market position than about a conviction that the field owes children and families a clearer answer to a deceptively simple question: is the therapy working?
The people who led the effort are, by and large, among the most recognizable clinical officers in the field. Joining Dixon were Dr. Adam Hahs, the Chief Clinical Officer of Caravel Autism Health; Dr. Hanna Rue of LEARN Behavioral; Jana Sarno, BCBA and Chief Clinical Officer of Hopebridge; and Dr. Kristine Rodriguez, Chief Clinical Officer of Autism Learning Partners, each the top clinical officer at their organization. Jeanine Weichelt, the Chief Clinical Officer of Helping Hands Family, completed the group.
Six of the seven spoke with Acuity for this article, willing, in other words, to be measured in the open; the seventh took part in the effort but is not featured here.
The CASP Treatment-Intensity Benchmark Seven ABA Providers Agreed to Be Judged Against
Last year, the Council of Autism Service Providers (CASP) published a white paper (a companion to the third edition of its practice guidelines) that did something the field had long circled without doing: it put numbers on what good early ABA should produce. Drawing on a new individual-participant meta-analysis led by the Norwegian researcher Dr. Sigmund Eldevik, it sorted comprehensive treatment into three intensity bands (roughly five to twelve hours a week, thirteen to twenty-five, and twenty-six to forty) and reported the gains a child might be expected to make at each: how many points on the Vineland, the standard measure of adaptive behavior, and what share should reach the non-clinical range.
The pattern was a dose-response. More hours, up to a point, tended to mean more progress. The question the seven set for themselves was deliberately modest, and Dixon keeps it that way. When different organizations apply the same benchmark to their own data, do they see similar results? “We are not trying to create a new standard for the
Dixon walked Acuity through the results in an interview. Across the group’s analysis of more than 5,000 children between the ages of three and seven, he said, the organizations’ real-world outcomes were broadly consistent with the published curve, with what he described as a “stair-step” pattern: a greater share of children in higher-intensity care had significantly greater improvements than in the lower-intensity groups. He also showed that the data cut against a familiar assumption, that large ABA organizations run every child through the same forty-hour prescription. The hours, he said, varied widely from child to child, which is what individualized care is supposed to look like.
That seven organizations who compete did this together is part of what makes it notable, and they were scrupulous about how. Dixon described a collaboration deliberately confined to clinical questions: no pricing, no contracting, no market strategy, only de-identified, aggregate outcomes measured against an outside benchmark.
Hahs, a former associate editor at the journal Behavior Analysis in Practice, framed the deeper significance as a matter of professional maturity. “For decades, the ABA industry has wrestled with defining what good looks like,” he told Acuity, acknowledging that organizations might have identified assessments and metrics that “paint their service in a positive light.” The seven, he said, came “hat in hand,” accepting that any shared methodology would carry compromises, in exchange for showing that the field can agree on the “goalposts of good”, and stick with them.
The appetite for that is, for lack of a better word, voracious. Sarno presented the multi-provider work at this year’s CASP conference, one of several ABA meetings where the group shared its findings and issued an open call for others to join, and told Acuity the reception was warm; attendees came up afterward asking how to launch something similar in their own organizations, which she read as a sign that peers are hungry for this kind of measurement. Before Hopebridge finalized its part, she ran the approach past the company’s Clinical Advisory Board, which she describes as a standing panel of clinicians, applied researchers, and policy and regulatory figures, and which she said helped sharpen the work. The most pointed note from the room was not doubt about the numbers but a call to invest in the people behind them: better training for behavior analysts in how these assessments are delivered and how cases are conceptualized. A benchmark, in the end, is only as good as the clinician entering the scores.
Why ABA Outcomes Data Is So Hard to Produce
Dixon’s forty to sixty hours were not a quirk of one organization’s systems. The software that runs ABA was built to document a service, submit a claim, and hold a single child’s record, not to analyze outcomes across a thousand children at once. Readiness for value-based care, the arrangement much of behavioral health says it wants, turns out to depend on closing that gap. It is “not only a question of whether providers deliver effective care,” Dixon said. “It is also a question of whether they can demonstrate that effectiveness reliably and at scale.” That readiness problem is one the whole field is only beginning to take on.
Rue, who once chaired the second phase of the National Standards Project, said the analysis worked because the participants settled the unglamorous questions first: a common definition of a direct treatment hour, the same outcome measure, the same collection intervals. Asked whether infrastructure is the real bottleneck to comparable data, she answered, “Largely, yes.” She pointed to efforts like ICHOM and an outcomes battery published by the Behavioral Health Center of Excellence as real progress, and named the piece still missing: consensus, across providers and trade associations, about what to measure in the first place.
The group concedes this cannot remain a capability of the largest organizations alone. Weichelt called the initial build the heavy lift and said she would still urge every provider, “large or small,” to make it. It is an aspiration with some friction behind it: the organizations best equipped to measure outcomes today are the larger ones, while many independents are already strained by fixed compliance costs that fall hardest where there is no scale. Part of what Rodriguez and others say they want is a framework light enough that a boutique practice and a multi-site organization can both actually use it.
What Counts as a Good ABA Outcome, and Why the Vineland-3 Is Only a Start
The measure at the center of all this is one the clinicians are the first to qualify. Hahs called the Vineland-3 a “level-set,” the tool each provider already used at set intervals in comparable ways, chosen because reaching for richer but dissimilar metrics “only further increases skepticism.” Caravel has its own real-time platform, PathTap™, and Hahs helped ICHOM define autism outcomes that reach well past adaptive behavior, but not every participant had those instruments. “We wanted progress at the known sacrifice of perfection,” he said, “where the latter is the antithesis of the former.”
Rodriguez, a co-author of that international ICHOM standard set, was emphatic about what the shorthand can miss. It is not the group’s finding, she told Acuity, that “more hours produce higher adaptive scores.” At ALP, the Vineland sits inside a broader battery that includes quality-of-life measures like the KIDSCREEN and the Child and Family Quality of Life scales. What caught her attention in her organization’s data was quieter: a relationship between how much a child needed at intake and the share of recommended hours they ultimately received, and so how many medically necessary skills there was time to build. She is careful, too, with the word that shadows intensive ABA. For a child who arrives profoundly affected, she said, the substantial investment of
Hahs is forthright about the objection such a project invites. An analysis run by providers on their own clients naturally raises the question of bias, and he names the threats himself: homogeneous sampling, repeated-measures effects, the pull toward flattering numbers. The contributors, he said, committed to “showing our clinical outcomes hands” anyway, on the view that acknowledged limitations are more useful than unexamined ones. He believes the work would hold up in peer review, while granting that it would need the usual scaffolding, preregistration and outside replication, to get there.
None of this sits in settled science, and the clinicians do not pretend otherwise. Whether more ABA reliably yields better outcomes is hotly contested: a 2024 meta-analysis in JAMA Pediatrics by Dr. Micheal Sandbank and colleagues found no clear link between the amount of intervention and children’s gains, and a line of research associated with the nonprofit Catalight has argued that outcomes can improve largely independent of hours. Hahs engages it directly rather than waving it off. A full response to Sandbank is “beyond the scope” of this first effort, he said, but he pointed to a subsequent analysis by Dr. Thomas Frazier and colleagues, who reworked the same data with IQ taken into account and concluded that dosage was associated with better outcomes after all. (Sandbank’s team has defended its original findings, and the exchange continues.)
Even Eldevik’s team, whose meta-analysis underlies the benchmark, warns against reading too much into it: the numbers apply to groups of children, not to any one child, and the studies behind them carry a real risk of bias, since children in the research were never randomly assigned to more or fewer hours. Dixon accepts the point and reframes it. If families gravitate toward more intensive programs in the research, they do the same in the clinic, which to his mind makes comparing an organization’s outcomes to the published benchmark a fair, like-for-like exercise rather than a flaw, and part of a larger case he has made for letting outcomes, not process, settle the field’s arguments.
No one in the group is inclined to stretch the result past what it can bear. Sarno, whose clinical roots are with the most severely affected children, was asked how well a sample of three-to-seven-year-olds speaks to older or more profoundly affected kids, where the evidence on intensity is thinner. She wouldn’t guess beyond her own data and pointed instead to the colleagues who serve those populations. It was a small act of restraint, and a revealing one: the benchmark travels only as far as the children it was built on, and the people using it seem to know it.
Who Defines ABA Quality, and What It Means for Payers and Access
The number of people with an autism diagnosis drawing on Medicaid or CHIP rose sixty-seven percent between 2021 and 2025, from 1.15 million to 1.92 million, and only a small share of them receive ABA at all. As states weigh caps on authorized hours and rewrite reimbursement, the clinicians are keenly aware that a benchmark can be read in more than one direction.
Weichelt hopes the data “should act to inform, but not dictate.” Dixon’s longer-standing concern is that if the field cannot show which care is good, others will define quality on its behalf. The group makes no claim that more is always better; several of them resist that reading. What the clinicians say they want is a conversation about quality that protects access rather than trims it, and they have company: CASP is building a national data platform so member organizations can benchmark against pooled numbers, a sign that the field is moving, from more than one direction at once, toward measuring itself.
Rodriguez’s version of success is a system in which payers reward quality through pay tied to a child’s quality of life, with frameworks light enough for a small practice and a large one alike. Meanwhile, Dixon does not expect a single moment of arrival; a year from now, he said, he would be glad to see more organizations measuring and sharing, payers asking sharper questions about clinical quality rather than only cost, and the field reaching past the Vineland toward measures of whether a life got better. He keeps returning to Thomas Kuhn: fields change, he said, not because someone outside announces a new standard, but because the people inside come to see the limits of the old way and show a better one. “We are trying to make it more normal for providers to measure results, share what they find, and be transparent about both the strengths and limitations of the data.”
Opinions and clinical expertise aside, seven organizations that had every commercial reason to keep their numbers to themselves decided, instead, to put them on the table. Whatever the data ultimately shows, that choice is the part the field is likely to remember.






