A single brain implant decoded speech and gesture at the same time, a first for people with paralysis
A UCSF team has shown a single brain implant decoding attempted speech and gestures at the same time, driving an on-screen avatar for people with paralysis. It is a genuine advance, and a small proof of concept, not a finished device or mind reading. Here is what it did, how well it worked, and what it is not.

Researchers at the University of California, San Francisco reported, in Nature Neuroscience on 14 September 2026, what they believe is the first brain-computer interface to decode intended speech and communicative gestures at the same time from a single implant. Three people with paralysis took part; two controlled a full-body avatar that gestured in near real time while their decoded speech appeared as on-screen text. Real-time accuracy for the larger vocabulary ran roughly 66 to 85 per cent across the two channels, on a small closed set of phrases and gestures. The standout scientific finding is that combined speech-plus-gesture is not simply the sum of the two done separately; it has its own neural pattern. This is an important proof of concept, not a take-home product, and it is not mind reading. Here is the careful version.
Most brain-computer interfaces that restore communication give a person back one channel: text, or a synthesised voice. Real conversation is not one channel. It is words carried along with nods, shrugs, a wave, a thumbs up, the timing cues that tell the other person when it is their turn. A team at UCSF has now shown, for the first time, that two of those channels can be pulled from the brain at once, through a single implant, with the gestures played back through an avatar and the words shown as text, in near real time. It is a real step forward. It is also small, invasive and early, and the most quietly significant part is not the avatar at all.
What the team actually built
The work, published in Nature Neuroscience on 14 September, comes from the lab of Edward Chang at UCSF, with bioengineering researcher Samantha Brosler as first author, and was funded by the US National Institutes of Health, through its deafness and communication-disorders institute. The hardware is a single high-density electrocorticography (ECoG) array, a thin sensor strip of around 253 electrodes laid on the surface of the sensorimotor cortex. Crucially, that sits on the brain rather than penetrating it, which is less invasive than the needle-like arrays some rival groups use, at the cost of some signal resolution. The signals run by wire to external computers, where machine-learning decoders translate them.
Three participants took part, living with paralysis from brainstem stroke or ALS. The headline result, simultaneous control of the avatar, was demonstrated in two of them. The vocabularies were deliberately small and closed: one participant worked with up to 10 phrases and 10 gestures, giving 100 possible combinations, another with five phrases and four gestures. The gestures were things like a wave, a nod, a handshake, a shrug or a thumbs up. And here is a detail that a lot of coverage blurred: the avatar performed the decoded gestures with its body, while the decoded speech was displayed as text on the screen. This is not primarily a talking face. The new thing is the body-language channel running alongside the words.
Why doing both at once is the hard part
You might assume that if you can decode speech, and you can decode a gesture, you can just run both decoders together. The team found it is not that simple, and that is the intellectually interesting result. The neural activity for speaking-while-gesturing is not merely the two isolated patterns added together; combined expression has its own signature in the cortex. Decoders trained only on isolated data struggled when the two happened at once.
The fix, and the finding, was to train the models on both isolated and simultaneous data. Doing so improved accuracy across the board and sharply reduced the rate at which one channel falsely triggered the other. The researchers also report that the models could handle some speech-and-gesture pairings not seen in training. The practical lesson for the field is that future multifunctional interfaces should probably be trained multimodally from the outset, rather than bolting separate single-purpose decoders together.
How well it worked, honestly
The accuracy numbers need to be read with their caveats attached. In the real-time simultaneous task for the larger, 100-combination vocabulary, the system decoded gestures at around 66 per cent and speech at around 70 per cent, rising to about 85 per cent for gestures and 75 per cent for speech in a conversational demonstration. The participant working with the smaller vocabulary did better still, reaching a median of 100 per cent for both channels across a few short conversation blocks. You will see that perfect figure quoted, but read it with care: it comes from a small number of blocks on a limited vocabulary, not from sustained, open-ended use. The fair summary is that the harder, larger-vocabulary task ran at roughly 66 to 85 per cent across the two channels, which is respectable for a first demonstration and nowhere near the reliability a daily communication device would need.
On speed, the researchers describe the system only as working in near real time. No specific latency figure was published, so anyone quoting milliseconds is inventing them.
Is it mind reading? What the brain implant does not do
Two clarifications matter, because the popular framing tends to run ahead of the science.
First, this is not mind reading. The interface decodes speech and movements the person is actively attempting to make. It does not read spontaneous private thoughts, and it should not be described as if it does. (A separate line of research elsewhere on decoding imagined "inner speech" is a different and more sensitive thing; do not conflate the two.)
Second, this is not a product. Three participants, two of them for the full demonstration, is a proof of concept, not a powered clinical trial. It requires open neurosurgery to place the array. It runs on a wired laboratory rig, not something you take home. Its vocabulary is a couple of dozen items at most (10 phrases and 10 gestures in the larger case), against the roughly 125,000-word vocabulary the best speech-only systems now handle. The team's own next step is to build a fully implantable, wireless version, and durability, natural-conversation performance and regulatory clearance are all still ahead. As Deb Lowe of the Stroke Association put it, the research "is in its early clinical stages, and further exploration is required to fully understand how it could truly benefit stroke survivors."
Is this Neuralink? How it compares
It helps to place this against the field, because different groups are optimising for different things. UC Davis reported a speech neuroprosthesis in 2024 that hit around 97.5 per cent accuracy over a 125,000-word vocabulary, the benchmark for raw speech decoding. Stanford's BrainGate researchers have pushed into decoding imagined speech. Neuralink and others are chasing high electrode counts with implanted, penetrating devices, while Synchron threads its sensor in through a blood vessel to avoid open-brain surgery altogether. UCSF's contribution here is not speed or vocabulary size, where others lead. It is breadth of function: two communicative channels, decoded together, for the first time. That is a different axis of progress, and arguably the one closest to how people actually communicate.
Chang framed the ambition plainly: "Conversation is about much more than the words being spoken. It's a multilayered, dynamic process involving the whole motor cortex. This proof-of-concept shows us it's possible for a BCI to restore some of this freedom and flexibility." Brosler was careful to bound the claim: "To our knowledge, this is the first demonstration that speech and communicative gestures can be decoded at the same time from a single brain implant in people with paralysis."
The people who stand to gain most, those with ALS, brainstem stroke or locked-in syndrome, are exactly the group for whom flattening communication down to a single channel costs the most, and also the group for whom consent and long-term data privacy raise the hardest questions. Restoring not just what someone says but some of how they say it is a meaningful expansion of what these devices can give back. The honest headline is that it works, in a lab, on a small scale, for the first time, and that the road from here to something a person could live with is long but now clearly pointed in a more human direction.
Frequently asked questions
What is a brain-computer interface?
A brain-computer interface (BCI) is a system that records brain activity and translates it, with software, into an action such as text, synthesised speech or control of a cursor or avatar. It reads signals of intended action, not private thoughts. See our explainer on how brain-computer interfaces work.
How does the implant help someone speak?
In conditions like ALS or a brainstem stroke, the speech-motor part of the brain can still fire, but the signal no longer reaches the muscles. Electrodes resting on the cortex capture that activity, and a decoder turns the attempted speech into words on a screen or a synthesised voice.
Can a brain implant read your mind?
No. It decodes attempted speech and movement from the motor cortex, within a small, trained vocabulary. It cannot pull out unspoken private thoughts, and describing it as "mind reading" overstates what it does.
Is this the same as Neuralink?
No. This is an electrocorticography (ECoG) grid resting on the surface of the brain. Neuralink uses fine electrodes that penetrate the tissue, and Synchron places its device inside a blood vessel. Different hardware, different trade-offs between signal detail and invasiveness.
How accurate is it?
On a constrained task with a small set of phrases and gestures, the larger-vocabulary participant ran at roughly 66 to 85 per cent across the two channels, and a participant with a smaller set reached a median of 100 per cent across a few short blocks. Read those as early research results, not everyday free conversation.
What is electrocorticography (ECoG)?
ECoG places a thin grid of electrodes directly on the surface of the brain to record activity. It is less invasive than electrodes that penetrate the tissue, at the cost of some signal resolution, and it is the method this study used.
Is the brain implant available to the public?
No. It is early-stage research in a small number of participants and it requires neurosurgery. There is no consumer product, no price and no promised timeline.
What was the actual breakthrough?
Doing both at once from a single implant. The team found the brain encodes combined speaking-and-gesturing differently from doing either alone, and training the decoder on that combination let it drive speech and gesture together in near real time.





