- Invasive brain implants allow speech attempts to be converted into voice almost in real time, restoring communication to patients with paralysis.
- Non-invasive techniques (MEG, fMRI, EEG) combined with large language and image models can reconstruct perceived or imagined phrases and scenes.
- Current advances manage to decode specific tasks and content, but they still do not allow us to read the entire spontaneous flow of the human mind.
- The expansion of these neurotechnologies demands strong ethical and legal frameworks to protect mental privacy and ensure responsible use.

The idea of reading minds with artificial intelligence has gone from a science fiction dream to a very serious scientific field, with laboratories around the world competing to decode what we think, see, or imagine based on brain activity. Today, we're no longer just talking about hypotheses: there are implants that convert thoughts into speech almost in real time and non-invasive systems capable of reconstructing sentences and even images from brain signals.
This scenario opens up a vast array of applications, from restoring communication to people with paralysis to developing new ways of interacting with computers. But it also raises major ethical dilemmas regarding mental privacy, consent, and potential misuse. Let's examine, calmly and in detail, what is really happening with AI-based neurotechnologies for thought decoding, what they can do today, and what challenges lie ahead.
Brain implants that turn thoughts into speech in seconds
One of the most spectacular advances in recent years is the development of a brain implant capable of transforming neural signals into audible speech in just a few seconds. Published in a top-tier scientific journal in 2025, this system represents a qualitative leap in brain-computer interfaces (BCIs) for people who have lost the ability to speak.
Until recently, most speech-oriented BCIs were rather clunky: the user had to mentally complete an entire sentence for the system to process it and generate speech or text output. This resulted in disjointed conversations, full of long pauses, far removed from the natural rhythm of everyday dialogue.
The new implant, however, has managed to reduce decoding time to about 3 seconds per word or group of words . It's not exactly as fast as a face-to-face conversation, but it comes much closer than previous technologies and allows for much smoother and more natural interactions.
The heart of this breakthrough lies in the combination of high-density electrodes and advanced AI algorithms . The implant, a flexible sheet with hundreds of contacts, is placed on the surface of the cerebral cortex responsible for speech control and records the combined activity of thousands of neurons simultaneously.
These signals are sent to an external system that, using deep learning models, identifies neural patterns associated with phonemes, words, and speech structures . Once decoded, they are sent to a speech synthesizer that generates audible speech in near real time.
Ann's case: regaining her voice after years of silence
Beyond engineering, one of the stories that best illustrates the human impact of these neurotechnologies is that of Ann, a woman who lost her speech after suffering a brainstem stroke in 2005. The damage left her in a situation where her mind remained intact, but her body had lost the ability to articulate words.
Eighteen years later, Ann underwent surgery in which a thin sheet with 253 electrodes was implanted on the surface of her speech-related cerebral cortex. This array simultaneously records the electrical activity of thousands of neurons that fire when she tries to speak, even though her muscles no longer respond.
The magic happens when this data is combined with artificial intelligence algorithms specifically trained for this case. For months, the system learned to associate specific neural patterns with words and articulatory movements that, in a healthy person, would produce sounds.
The researchers wanted to go a step further and weren't satisfied with a generic robotic voice. They retrieved video recordings of Ann's wedding , from before her stroke, and used them to train the voice synthesizer. As a result, the system's audio output closely resembles how her real voice sounded.
The result is that Ann not only regains the ability to communicate, but she does so with a voice she recognizes as her own , something that has an enormous emotional impact. In this case, technology has not only restored a function, but also a part of her identity.
How a speech decoding implant actually works
From a technical point of view, the process that allows thoughts to be converted into speech with an implant can be divided into several stages, all of them closely coordinated and supported by artificial intelligence models.
First, when a person attempts to pronounce a word or phrase , even if they cannot move their mouth muscles, complex electrical patterns are generated in their cerebral cortex. The implant's electrodes record these voltage variations with very high spatial and temporal resolution.
Next, this raw signal passes through a data processing system that cleans it and transforms it into a set of numerical features that AI algorithms can handle. The system must separate the relevant information from background noise , which includes non-speech-related neural activity and other interference.
Once these features are extracted, deep neural networks come into play. These networks have been previously trained through long sessions in which the user attempts to produce known words while the system records the associated brain signals . After many repetitions, the AI learns to map these patterns to specific linguistic units.
Once the model is calibrated, it can, during everyday use, predict which word or sequence of phonemes the person is trying to produce simply by analyzing the activity recorded in real time. That sequence is then sent to a speech synthesizer, which generates sound with very low latency, just a few seconds.
The major improvement over previous systems is that it's no longer necessary to wait for the user to finish a full sentence to generate voice output. The model is able to continuously update its prediction as new neural data arrives, allowing for a much smoother conversation flow.
Clinical applications and future projections of invasive BCIs
The ability to transform thoughts into speech almost in real time has enormous implications for people with speech disabilities caused by ALS, stroke, high spinal cord injuries, or other neurological conditions. For many of them, regaining a rapid form of expression can completely change their daily lives.
A brain-computer interface that functions with this fluidity not only allows for yes and no answers, but also the expression of nuances, emotions, and complex opinions . This has a positive impact on the quality of life, autonomy, and psychological well-being of both patients and their families.
Beyond speech, researchers are considering extending these decoding methods to other cognitive functions affected by injury , such as motor planning or certain aspects of memory. In theory, if the neural patterns involved can be recorded and modeled, personalized assistance or rehabilitation systems could be designed.
In the field of basic neuroscience, these tools open a unique window into how the brain represents language, communicative intentions, and semantics on a large scale. Being able to decode in real time what a brain is trying to say provides data that was previously impossible to obtain with such precision.
However, all this potential comes with significant challenges: delicate surgeries, possible long-term side effects, and still very high costs . Therefore, alongside the development of implants, a second line of research has gained considerable momentum: non-invasive neurotechnologies for reading thoughts.
Non-invasive neurotechnologies: reading thoughts without surgery
In recent years, striking results have been published from groups that, without opening the skull, have managed to reconstruct continuous sentences and descriptions of scenes from brain activity measured with techniques such as functional magnetic resonance imaging (fMRI), magnetoencephalography (MEG) or EEG.
One of the most talked-about developments is the so-called "mental subtitling" achieved by a team of Japanese neuroscientists. Using fMRI and advanced language models, they have created a system capable of generating texts that describe what a person is seeing, imagining, or hearing with a surprising level of detail.
This system not only produces sentences about external stimuli (for example, a video the subject is watching), but can also capture aspects of how the brain internally represents the world before those contents are formulated as conscious words.
The researchers trained the AI using more than 2.000 videos with their associated subtitles , converting each one into a numerical "signature of meaning." Simultaneously, they measured the brain activity of several volunteers using fMRI while they watched these same videos. In this way, the model learned to associate specific brain patterns with particular semantic signatures.
Then, when the system receives a new image of brain activity in response to a video it has not seen, it is able to predict which meaning signature is most likely and, from there, generate sentences with another language model that closely approximate what the person is perceiving or imagining.
Another major leap forward comes from groups like Meta AI and the Chinese Academy of Sciences, who have shown that it is possible to reconstruct continuous sentences from MEG , a technique that records the tiny magnetic fields generated by neuronal activity without the need for implanted electrodes.
In this case, the key has been combining MEG with Transformer-type language models , similar to those used by the most advanced AI chats. Instead of focusing on identifying individual phonemes, the systems analyze global magnetic patterns and let the language model "complete" the most probable sentence based on the context.
From noise to meaning: how AI cleans and understands brain signals
Traditionally, one of the problems with non-invasive techniques like MEG, fMRI, or EEG is that the useful signal is buried in a great deal of noise . Brain activity is extremely complex, and what is of interest for decoding language or images is only a small part of everything that is happening at once.
In MEG, for example, when we think of a word or hear a phrase, thousands of neurons fire in sync , generating minuscule magnetic fields. Detecting them from outside the skull is like trying to hear a whisper in a stadium full of screaming people.
The current strategy to overcome this obstacle is based on the semantic reconstruction of speech . Instead of trying to translate each signal fragment into a specific letter or sound, AI models learn to associate complex dynamic patterns with meanings at the level of words, phrases, and even scenes.
Architectures like Transformers allow the system to take context into account: if a certain sequence of brain patterns indicates that the person is hearing or imagining a story , the language model can fill in the less clear parts based on the probability of certain words appearing.
Research led by Chinese teams has also shown that the brain organizes language hierarchically. There is one layer that reflects the intention to communicate and another that captures the specific content of the message. Identifying these "layers" allows AI to separate what is mere background noise from what is actually part of the linguistic process.
Models like BP-GPT combine functional magnetic resonance imaging (fMRI) data, which offers high spatial accuracy but is slow, with MEG, which is very fast but less spatially detailed . fMRI "teaches" the model where to look, and then MEG provides a rapid snapshot of how language evolves over time, improving the ability to discriminate between heard and imagined speech.
From image to text: MinD-Vis systems and visual decoders
In addition to language, AI is making remarkable progress in reconstructing images from brain activity . A good example is MinD-Vis, a system designed to translate patterns obtained with fMRI into images that closely resemble what the subject is seeing.
The process is again divided into an encoder and a decoder. The encoder uses convolutional neural networks to mimic the brain's visual processing stages and translate the input images into a feature space that can be associated with the recorded brain signal.
The decoder does the reverse: it receives the brain activity pattern and, using diffusion-based generative models , reconstructs a high-resolution image that closely resembles what the person was actually seeing.
In recent work, researchers at Radboud University have improved these decoders by incorporating attentional mechanisms that allow them to focus on particularly informative brain regions during reconstruction. As a result, the generated images are even more precise and detailed.
Although these reconstructions are neither perfect nor photographic, they show that the correspondence between brain patterns and visual information is robust enough for AI to successfully exploit it, which in the medium term could serve as a basis for visual aids, diagnoses, or new forms of art and communication.
DeWave and other EEG systems that translate silent thoughts
On the EEG front, which measures the brain's electrical activity with electrodes placed on the scalp, proposals like DeWave stand out—a non-invasive system that translates silent thoughts into text . Here, bulky machines and operating rooms are no longer needed; instead, a cap with sensors is sufficient.
DeWave works by recording the EEG signal while the person silently reads sentences or thinks about specific words . From large volumes of data, deep learning models detect patterns in brain waves that correlate with specific linguistic meanings.
The system introduces a technique called discrete coding , which transforms EEG segments into unique numerical codes organized in its own "codebook." Each code is mapped to nearby words or linguistic fragments in that space, allowing sentences to be constructed.
In practice, DeWave also uses an encoder-decoder scheme. The encoder, based on BERT (a bidirectional language model), converts the EEG signal into symbolic representations , while the decoder, of the GPT type, transforms those symbols into written words.
The results don't yet allow for a fluid, real-time conversation, but it's already possible to grasp the general meaning of complete sentences and many of their keywords . There are still grammatical errors and some lack of fluency, but the "technical bridge" between brainwaves and written text has been built.
Can we say that mind reading is already possible?
With all these advances on the table, it's tempting to claim that mind reading is already an everyday reality , but the current situation is much more nuanced. The technologies available today shine when it comes to very specific, well-trained tasks in controlled environments.
In other words, we can decode certain types of thought or perception with very good accuracy (for example, listening to a story, seeing a specific image, or trying to pronounce words) when the system has been specifically calibrated for that person and that set of stimuli.
However, we are still far from an AI capable of reading the continuous and spontaneous flow of the human mind in all its diversity: memories, abstract thoughts, subtle emotions, daydreams, dreams… The challenge is that mental states are extremely rich and dynamic, and their reflection in the brain does not follow a simple dictionary.
Even so, in tasks such as cursor control, word prediction in a narrative, or visual scene reconstruction, accuracy rates have improved dramatically in just a few years. The development of large-scale language and vision models has been a key catalyst for this progress.
It's reasonable to assume that, with better sensors, more data, and even more powerful models, the accuracy and speed of these decodings will continue to increase . The big question is how the ethical implications of that capability will be managed.
Mental privacy, consent, and ethical risks
The possibility of inferring what a person sees, imagines, or tries to say will inevitably raise intense debates about the privacy of thoughts . If the mind ceases to be a completely inaccessible space, it will be necessary to clearly define what can be decoded, when, and under what conditions.
Currently, all these systems require active user participation and extensive calibration sessions . You can't "spy" on someone remotely with a MEG headset without them knowing, much less read their thoughts without a specific experimental environment.
Even so, the researchers themselves insist on the need to develop legal and ethical frameworks that protect neural information as a particularly sensitive type of data. Brain signals can reveal aspects of health, preferences, moods, and internal processes that a person might not want to share.
There is also a risk of misunderstandings: even the best models make mistakes, which can lead to misinterpretations of neural signals . In clinical, legal, or occupational settings, this could be especially serious if not handled with caution.
Therefore, there is a need for clear policies on informed consent, transparency in the use of data, and strong privacy protection , so that technology is geared towards empowering users and not monitoring them.
In parallel, there is a more philosophical debate about the extent to which the possibility of externalizing thoughts and perceptions can change how we understand intimacy, identity, and communication . This is new territory that will require joint reflection from scientists, legal experts, philosophers, and society at large.
The current state of AI-powered neurotechnologies for thought decoding paints a picture in which, on the one hand, patients who had lost all ability to speak recover their voice thanks to brain implants that translate neural signals into synthetic speech almost instantly, and on the other hand, external sensors such as MEG, fMRI, or EEG allow the reconstruction of phrases and images without the need for surgery, relying on large-scale language and vision models; although there are still significant limitations in speed, overall accuracy, and generalization to complex thoughts, the combination of advanced AI, new sensors, and appropriate ethical regulation points to a future in which the gap between brain activity and communication becomes increasingly smaller, provided that the safety, consent, and privacy of those who benefit from these technologies are prioritized.
