Fifty-one behavior analysts rated AI and expert clinician responses blind. They preferred the AI, agreed with it more, and could not identify the source.
Key Takeaways
- The finding is preference, not accuracy: Participants preferred ChatGPT-4 responses to expert clinician responses and agreed with them more, in a blind comparison. The study did not establish that the AI answers were clinically better.
- They could not tell the difference: Participants were unable to reliably distinguish AI responses from human ones, which is the result with the most direct operational consequences.
- Reported use is far below reported exposure: Fewer than one in six participants said they use AI in their own behavior-analytic work, in a field where vendors now ship AI documentation by default.
- The version circulating is not the study: Retellings have described graduate students rating a professor’s answers. The published paper describes certified behavior analysts rating an expert clinical team.
The paper is Peck, O’Brien, Bourret, and Agostinelli in the Journal of Applied Behavior Analysis, published online in August 2025 and running in volume 58. Researchers at Western New England University and the New England Center for Children posed ABA-specific questions to ChatGPT-4 and to an expert clinical team, then asked 51 behavior analysts to rate the responses without knowing which was which. Participants significantly preferred the AI responses and agreed with them more. In the next phase, asked to identify the source, they could not do it reliably. Asked finally about their own practice, fewer than one in six reported using AI in their behavior-analytic work.
It is a small study with a narrow question, and it has become the most cited piece of AI research in this field, which means it is also the most misremembered.
What the JABA Study Did Not Test
Preference is not accuracy, and the paper does not claim otherwise. Nothing in the design establishes that the AI answers were clinically better, only that behavior analysts liked them more and agreed with them more. There are well-documented reasons a generated response wins a preference rating without being more correct: it is typically longer, more organized, more thoroughly hedged, and written in the register of a confident expert, all of which read as authority on a rating scale.
The other limits are ordinary ones. Fifty-one participants is a modest sample. The questions were posed by researchers rather than drawn from live caseloads, so the comparison happens outside the context where clinical judgment does its work. And the model tested was ChatGPT-4, now well behind what practitioners carry on their phones, which makes the finding a floor rather than a ceiling. Anyone citing this paper as proof that AI outperforms clinicians is citing a study that was not run.
What the Study Turns Into When It Travels
From a conference stage in Boston this month, Adam Ventura raised the paper as evidence that the field needs to rethink what it asks people to produce. His account of it: researchers took a small group of graduate students, gave them a comparative analysis of answers to behavior-analytic questions, and the graduate students reliably picked the AI-generated answers over the professor-generated ones. He was careful to hedge his memory of it, telling the room he believed it came out the previous summer and that they could look it up.
The drift is small and it changes the stakes completely. Graduate students preferring machine answers to a professor’s is a story about students. Certified practitioners preferring machine answers to an expert clinical team’s, and being unable to tell them apart, is a story about the people currently writing treatment plans. The second version is the published one.
What a Preference Rating Cannot Catch
Ellie Kazemi, who moderated, raised a limit the study does not reach. Her work developing clinical simulations has left her skeptical that these systems understand human interaction at all: what they return, she argued, is assembled from what already exists on the internet and is void of context and of the nuances of how people actually speak to one another. Cultural difference is where she sees the technology continuing to struggle, and it is precisely what a rating scale for a written answer will not detect.
Dan Dube, the panel’s technologist rather than its clinician, brought a market observation that cuts against alarm. He described watching large employers announce that AI would let them reduce headcount, then reading about productivity and quality falling afterward. His conclusion was that the evidence so far shows these systems cannot effectively replace people and work best as complements. That is compatible with the study rather than contrary to it: a tool can produce output that experts prefer and still be unable to hold a case.
Why Preference Without Discrimination Is the Operational Problem
For an operator, the finding that matters is not that the machine wrote something good. It is that trained clinicians could not tell. Any review workflow that depends on a supervisor noticing that a document reads like a machine wrote it is built on a discrimination the study suggests people cannot make. That has direct consequences for documentation, where federal auditors have made note quality the enforcement frontier. It also complicates the review architecture the field has been building, from the attestation workflows Massachusetts clinicians described this spring to vendor integrations arriving faster than anyone can evaluate them, including an AI assessment integration two companies say is already live without disclosing terms. Every one of those designs assumes a reviewer who can tell when something is off.
The gap between 16 percent reported use and the volume of AI now embedded in practice management platforms is its own finding. Either behavior analysts are using these systems without counting it as use, which is what happens when a feature is built into the software rather than opened in a browser tab, or the study caught the field at a moment that has already passed. Both readings point the same way: nobody knows how much AI-drafted material is currently moving through clinical documentation, including the organizations producing it.
What the Panel Proposed Instead of Alarm
Ventura’s response to his own reading of the research was not restriction. It was that the field is measuring the wrong output. Education has historically trained people to produce correct answers, he argued, and the systems now produce correct-sounding answers on demand, so the valuable skill has shifted to asking. His framing was behavior-analytic: prompting is manding and pinpointing, identifying what you want and asking for it precisely, which he claimed behavior analysts are better positioned to teach than anyone.
Kazemi described her own use as a feedback loop rather than a drafting tool, feeding in her work and asking the model to respond as a clinician would, then as a reviewer would, which she said is faster and more candid than asking friends who like her. She also pointed to work by Claire St. Peter on translating dense clinical documents into task lists for teachers, where implementation reportedly improved because the format matched the reader. That is a narrower claim than most vendor marketing makes, and closer to what behavior analysts building caregiver-facing tools are attempting: translation for a specific reader, not judgment.
The counterweight came from Shelby Dorsey, whose six-year-old spent a day deciding that a card trick on YouTube was AI and that a wrong answer from a smart speaker was authoritative, and arrived at the useful conclusion that a computer should never be trusted completely. The adult version of that lesson is what the study measures, and it does not arrive on its own.
What Would Actually Settle the Question
The study is a starting point nobody has built on. The version worth running uses questions drawn from real caseloads rather than composed for the experiment, scores responses against clinical criteria rather than preference, and tests current models rather than one that shipped three years ago. It would also measure what operators most need to know, which is whether human review catches errors in generated clinical text at a rate better than chance.
Until someone runs it, the honest position is narrow. In one blind comparison, certified practitioners preferred machine answers and could not identify them. That is enough to justify caution about any workflow whose safeguard is a supervisor noticing, and not enough to justify anything else. It is also a reminder that this field is buying faster than it is studying, in a market where the people building these products move considerably quicker than the literature underneath them.





