bci-simultaneous-speech-gesture-decoding
A single cortical surface implant decoded speech and gestures simultaneously in people with paralysis. Decoders trained separately on each task failed during combined use; training on joint data solved the problem.
A team led by neurosurgeon Edward Chang at the University of California, San Francisco implanted a single 253-electrode array on the brain surface in three people with severe paralysis (two after brainstem stroke, one with amyotrophic lateral sclerosis, a disease that progressively disables muscle control). The device simultaneously read out attempted speech and attempted gestures, and the decoded signals drove a personalized virtual avatar of each participant in real time. The results were published on September 14 in Nature Neuroscience.
Previously, brain-computer interfaces decoded speech or body movements only in isolation: in earlier work by other groups, a simultaneous attempt to speak corrupted cursor movement decoding, and vice versa. The reason lies in the cortex itself. The regions responsible for arm movements and for speech partially overlap, and when both are performed at the same time, the pattern of neural activity in those regions changes. A model trained only on isolated attempts did not capture this shift.
Chang's team tested this directly: separately trained speech and gesture decoders, when faced with simultaneous attempts, more often misclassified the signal as rest. The fraction of missed attempts in one participant reached 35%. The solution came not from new electrodes but from the composition of training data: models trained on both isolated and simultaneous examples maintained high accuracy in both modes. They preserved their ability to distinguish speech from gesture by relying more heavily on electrodes specific to a single modality. Each decoder was additionally taught to recognize the "other" activity as rest: the speech model learned to stay silent during gestures, the gesture model learned to stay silent during speech, and false activations during parallel operation nearly disappeared. On novel combinations of phrase and gesture, the model in one participant performed as well as on previously seen pairs: it generalized the rule rather than memorizing specific pairs.
The decoded signals drove a full-body avatar: gestures animated the body, speech appeared as text. In a simulated conversation, one participant achieved 100% accuracy for both speech and gestures; another reached 75 to 85%. A third participant withdrew from the trial, and the demonstration rests on data from the remaining two. The work continues a line from the same laboratory: in 2023, a similar implant already controlled an avatar through speech and facial expressions in one patient, but the facial expressions there were part of the speech articulation itself. The new result adds an independent channel of body movement and tests it simultaneously with speech.
Conversation is far more than the words spoken. It is a multilayered, dynamic process that engages the entire motor cortexsays Chang. Neurologist Daniel Rubin of Massachusetts General Hospital, who was not involved in the study, called the work "an important step toward developing truly natural systems for restoring function in people with paralysis."
The vocabulary remains limited: 10 phrases and 10 gestures. Frozen decoders without retraining maintained accuracy above chance level for days and months afterward. The principle found here reaches beyond the current vocabulary: it is a step toward neural interfaces that simultaneously read out multiple coordinated channels of human intent.