- Mar 20
Reading between the lines: my perspective on sham neurofeedback
- Brendan Parsons, Ph.D., BCN
- Neurofeedback, Neuroscience, Practical guide
There is no graceful way to say this, so I may as well just say it plainly: this is a blog post about my own paper. Which is either delightfully transparent or faintly embarrassing, depending on your tolerance for academic self-reference. I am choosing to believe it is the former.
But I also think this article deserves that kind of direct treatment. It is not just another summary of a new publication. It is, for me, a push towards a tipping point about how we evaluate neurofeedback, what kinds of evidence we keep privileging, and whether the field is finally ready to outgrow methodological habits (and unfavourable bias) borrowed from interventions that do not work the way neurofeedback works.
So rather than pretend to stand at a comfortable distance from the paper, I want to do something more honest here: tell the behind-the-article story. Why I wrote it. What problem I think the field keeps missing in the literature even while enacting it every day in the clinic. Why I believe the sham-control conversation has been framed too narrowly for too long. And why I think this matters not just for research design, but for the future credibility of neurofeedback itself.
Why I wrote this paper
For years, I kept running into the same pattern in neurofeedback discussions. Someone would raise a fair question about evidence, someone else would point to sham-controlled trials, and the conversation would quickly collapse into a familiar binary: either sham proves neurofeedback works, or sham proves it does not. What bothered me was not that people were asking for rigor. They should. What bothered me was the assumption underneath the whole exchange: that sham neurofeedback is equivalent to an inert placebo and therefore represents the cleanest possible test.
The more I looked at that assumption, the less convincing it became.
Neurofeedback is not a pill. It is not a static ingredient delivered to a passive recipient. It is an interactive training process. People engage with the feedback. They try things. They notice internal shifts. They repeat patterns that seem to work. Clinicians adjust thresholds, shape difficulty, modify protocols, and help participants stabilize useful states. Once you really take that seriously, the idea of a neatly “inactive” sham starts to wobble.
That wobble is what led to the paper. I wanted to examine whether the field had been treating sham as a gold-standard control without adequately asking whether it was actually inert, or even whether inertness was the right benchmark in the first place.
The core argument, in plain English
The heart of the article is simple: sham neurofeedback often is not sham in the way people imagine.
Even when the feedback stream is not driven by the participant’s real-time EEG, the experience still contains many elements that may be behaviorally active. Participants still receive structured audiovisual feedback. They still invest effort. They still engage expectancy, attention, and self-regulatory attempts. They may still experience a sense of agency. They still interact with a therapist or researcher. And, depending on how the thresholds and feedback schedules are implemented, the so-called sham may still overlap with genuine target states often enough to preserve a meaningful portion of the learning process.
That last point was especially important to me. A sham condition can look non-contingent on paper while still functioning as partially contingent in practice. If the feedback schedule rewards a participant around 80 percent of the time, and the participant is spending meaningful amounts of time in the intended state anyway, then a surprising amount of the feedback they receive may still operate as effective reinforcement. In the paper, I used an illustrative model showing how such designs may preserve a substantial proportion of functionally contingent events.
That does not mean sham and active neurofeedback are the same thing. They are not. But it does mean sham should be interpreted as much closer to a partially active comparator than to a true placebo. And that changes how we should interpret both positive and null results.
The part I think the field has underestimated
I suspect the field has sometimes been so eager to look scientifically respectable that it has accepted a version of rigor that does not actually fit the phenomenon being studied.
I understand why. Neurofeedback has spent a long time under suspicion. When a field feels scrutinized, it becomes very tempting to adopt the most familiar badge of legitimacy available, and in clinical science that badge is often the double-blind, sham-controlled trial. The problem is that not every intervention can be cleanly evaluated with pharmacology logic.
In medication research, the active ingredient can, at least in principle, be subtracted. In neurofeedback, the intervention is made of contingency, learning, salience, repetition, expectation, adaptation, and guided self-regulation. Those are not side effects surrounding the treatment. They are part of the treatment. So when a control condition preserves many of those ingredients, the comparison is no longer active treatment versus inert placebo. It becomes something more complicated: one learning environment versus another learning environment with overlapping active features.
That is not a reason to lower standards. It is a reason to become more precise about what question a study is actually asking.
Why double-blinding is not automatically neutral here
One of the more uncomfortable points in the paper is that, in clinical efficacy trials, strict double-blinding may distort neurofeedback delivery rather than simply purify it.
That statement is easy to misread, so let me be careful with it. I am not arguing against blinding as such. Assessor blinding is valuable. Role separation is valuable. Pre-registration is valuable. Transparent reporting is valuable. What I am questioning is the assumption that full double-blinding is always the highest form of rigor, even when it prevents the intervention from being delivered in the adaptive way it is ordinarily meant to be delivered.
In real neurofeedback practice, clinicians do not usually behave like pharmacists handing over identical capsules. They observe. They calibrate. They adjust thresholds. They change parameters based on response patterns, tolerability, and engagement. They help clients recognize state shifts and refine strategies for reproducing them. If blinding removes or restricts that layer of adaptation, then we may no longer be studying neurofeedback as it actually functions in the real world.
That is a big deal. Because once the treatment has been stripped of a meaningful part of its mechanism, a null result becomes harder to interpret. Did neurofeedback fail? Or did the study test a constrained version of it that suppressed some of the very processes through which it typically works?
What I am not saying
Whenever I make this argument, I feel a strong need to draw boundaries around it, because I know how easily it could be caricatured.
I am not saying sham has no value. It can be useful in mechanistic studies, especially where the goal is to isolate target engagement or examine specific causal pathways. I am not saying every sham protocol preserves the same degree of contingency. I am not saying all positive clinical outcomes in neurofeedback can be attributed to the mechanisms I emphasize. And I am certainly not saying the field should relax methodological standards and replace them with vibes, intuition, and confident anecdotes delivered over a nice spectral display.
What I am saying is narrower and, I think, more defensible: the meaning of sham depends on the study aim, the implementation details, and the actual ingredients preserved in the control condition. If researchers do not measure those ingredients, or at least report them transparently, then the interpretation of the study becomes shakier than it first appears.
Why this matters for credibility
This may be the part I care about most.
I do not think neurofeedback becomes more credible by forcing itself into a research design that makes intuitive sense for drugs but conceptual nonsense for learning-based interventions. I think it becomes more credible by being explicit about what it is, what it is not, and how its mechanisms should be tested.
That means distinguishing clinical efficacy questions from mechanistic questions. It means reporting thresholding rules in enough detail that another researcher can tell whether the sham condition was truly non-contingent or merely nominally so. It means actually recording and measuring residual contingency in the control group rather than treating it as a theoretical concern or assuming it away. It means measuring expectancy, engagement, and perceived agency instead of pretending they are just annoying contaminants. It means being honest that adaptation is sometimes part of treatment rather than a nuisance variable to be scrubbed away.
Most of all, it means accepting that rigor is not a costume. It is not something we achieve by copying the outer form of a respected trial design. Rigor comes from matching method to mechanism. That, to me, is the deeper issue beneath the paper.
The behind-the-article feeling
If I step outside the argument for a moment, I can say this more personally: I wrote this paper because I felt the field needed a cleaner vocabulary for a messy problem.
Too many conversations about neurofeedback get stuck in old scripts. Advocates can become defensive. Critics can become dismissive. Researchers can end up arguing over whether a sham trial was “good enough” without asking the prior question of what that sham condition was really doing. I wanted to interrupt that script.
I also wanted to write something that clinicians would recognize as real. Many practitioners already know, at a practical level, that neurofeedback is not just a signal-delivery device. It is a shaped, relational, adaptive process. Yet the research language we often use to defend or criticize it can flatten that reality beyond recognition. I do not think the solution is to abandon rigor in favor of clinical intuition. I think the solution is to build better bridges between the complexity of practice and the precision of research design.
In that sense, this article felt important to write because it was not just about sham controls. It was about methodological fit. And methodological fit, in my view, is one of the most underappreciated determinants of whether a field matures.
Where I hope this goes next
My hope is not that everyone suddenly agrees with me. (Well... maybe at least on some level.)
My hope is that the paper sharpens the conversation.
I want researchers to become more explicit about whether they are studying symptom change, target engagement, mediation, or comparative effectiveness. I want sham conditions to be described with enough granularity that we can estimate residual contingency (when it is not explicitly measured) rather than casually assuming it away. I want the field to think harder about assessor blinding, role separation, transparent adaptation rules, and more aim-matched comparators. And I want clinicians and researchers to stop talking past one another as though ecological validity and methodological rigor were mortal enemies.
If the paper helps with any of that, then it will have done something worthwhile.
And yes, I realize there is something slightly strange about writing a blog post that says, in effect, “here is why my own paper matters.” But perhaps that is also fitting. The article itself is an argument for greater conceptual honesty in neurofeedback research. It would be a bit silly to discuss it from a fake distance.
So this is the honest version: I wrote the paper because I think the sham-control issue has been misunderstood in ways that ripple outward into study design, interpretation, and field-wide credibility. I think neurofeedback deserves tougher thinking, not softer standards. And I think the way forward is not to imitate pharmacology more convincingly, but to study learning-based brain training on terms that actually reflect what it is.
Conclusion
If I had to reduce the whole article to one sentence, it would be this: sham neurofeedback is often treated as though it removes the active ingredient, when in reality it may preserve more of that ingredient than researchers realize.
That is why I wrote the paper, and that is why I think it matters. Not because it lets neurofeedback off the evidentiary hook, but because it asks the field to become more exacting about the hooks it uses. Better control design will not weaken neurofeedback research. It will strengthen it. And if that conversation finally starts moving in a more sophisticated direction, then this paper will have done exactly what I hoped it would do.
Reference
Parsons, B. (2026). Rethinking control conditions in clinical neurofeedback trials. Discover Neuroscience, 21, Article 21. https://doi.org/10.1186/s13064-026-00252-x