The Brief
Health & Medicine 5 min read

Speaking and Gesturing at Once: What 88% Accuracy Means

NAVION

Share

Human communication is rarely just words. A nod, a wave, a shake of the head: these gestures carry meaning that speech alone cannot fully convey. For people living with paralysis who have lost the ability to both speak and move, that layered quality of communication has been doubly out of reach. A research team at the University of California, San Francisco has now taken a concrete step toward restoring it, using a brain implant to simultaneously decode intended speech and gesture from neural activity.

The Problem With Doing Two Things at Once

Previous experimental brain implants have shown promise in restoring either speech or movement to people with paralysis, but combining both has proven difficult. The reason is anatomical. To capture the signals needed for both functions, an implant must cover a large portion of the sensorimotor cortex, the brain region involved in controlling speech and movement. That is a significant engineering and surgical challenge.

The UCSF team addressed this by testing an iPhone-sized implant placed directly in that region of the brain. Two participants took part in the study. The first, referred to as Bravo-1r, had become paralysed following a stroke. The second, Bravo-6, had amyotrophic lateral sclerosis, the most common form of motor neuron disease. Both had lost nearly all limb movement and were unable to produce intelligible speech.

The experimental setup was methodical. Participants repeatedly attempted to produce ten phrases, including greetings like “hello” and “how’s it going?”, and ten gestures, such as waving, clapping, and shaking their head. Sometimes they attempted speech and gesture separately; other times, simultaneously. The brain activity recorded during these attempts was used to train customised machine learning models for each participant, models built to predict what the person intended to say or do.

What the Numbers Actually Show

The results were tested by asking participants to attempt phrase-and-gesture combinations the models had not previously seen. Predictions from the models drove an avatar that resembled each participant: the avatar moved to reflect the intended gesture, while the intended speech appeared as text on screen.

The accuracy figures are worth examining carefully. For Bravo-1r, gesture prediction reached 88 per cent accuracy and speech prediction reached 84 per cent. For Bravo-6, the figures were 66 per cent for gesture and 70 per cent for speech. As Samantha Brosler, the researcher leading the work, noted, a model operating on pure chance would achieve roughly 9 per cent accuracy given the range of options available. The gap between chance and these results is substantial.

Henri Lorach at the University of Lausanne offered a measured assessment: the work is robust, but the vocabulary of phrases and gestures remains limited, and accuracy would need to be higher for practical daily use without frustration. That is an honest framing of where the technology stands. Impressive in a research context; not yet ready to replace the full complexity of human conversation.

The team is continuing to refine the models and adjust the underlying algorithms. A future direction Brosler identified is moving beyond avatars entirely: if a participant’s joints and muscles are in adequate condition, the system could eventually drive movement in the person’s own body and generate a synthetic voice, rather than routing everything through a digital representation.

Why This Matters Beyond the Lab

Here is what most coverage of brain-computer interface research tends to underemphasise. The technical achievement, decoding two simultaneous streams of intended communication from neural signals, is significant. But the deeper point is about what communication actually is.

As Brosler put it, saying “maybe” while nodding carries a very different meaning than saying “maybe” while shaking your head. That distinction is not a minor nuance. It is the difference between agreement and doubt, between sincerity and irony, between a person being understood and a person being misread. For someone who has lost the ability to speak and move, the inability to layer those signals together is not just a physical limitation. It is a barrier to being fully present in a conversation.

Brain-computer interfaces are often discussed in terms of restoring function, and that framing is accurate. But this research points toward something more specific: restoring the texture of communication. The goal is not simply to give someone a way to produce words. It is to give them back the ability to mean things in the full, embodied sense that human interaction depends on.

The use of machine learning here is not incidental. The models are trained on each individual’s unique neural patterns, which means the system is personalised in a way that generic assistive technology cannot be. That personalisation is also a constraint: it requires substantial recording and training time before the system becomes useful. Scaling this approach, and making it accessible rather than experimental, remains a long road.

In Short

A brain implant tested in two people with paralysis successfully decoded intended speech and gesture simultaneously, achieving accuracy well above chance. The system uses personalised machine learning models trained on each participant’s neural activity, and currently operates through an avatar. Accuracy ranges from 66 to 88 per cent depending on the participant and the task. Researchers acknowledge that higher accuracy and a broader vocabulary will be needed before the technology can serve as a practical communication tool. The significance of the work lies not just in the technical result, but in what it targets: the layered, multimodal nature of human communication that words alone cannot capture.

Based on reporting from New Scientist.

Written by

NAVION