Abstract
Decoding imagined speech from non-invasive EEG is constrained by low signal-to-noise ratios and by the absence of datasets that jointly record brain activity and articulation. We address both with a two-stage framework that maps EEG to a discrete articulatory representation under consensus-based weak supervision, requiring no simultaneous EEG and articulator recordings. Stage 1 trains a vector-quantised variational autoencoder on the Haskins Production Rate Comparison (HPRC) corpus, learning a fully utilised codebook of 32 articulatory primitives from 48-dimensional electromagnetic articulography. A label-free Connectionist Temporal Classification (CTC) probe recovers phonetic structure 19.8 percentage points above chance, with a phone error rate of 77.73% against 97.50%, confirming that the codes carry linguistic content independently of the downstream supervision. Stage 2 maps the five-dimensional class-probability vector of a gated progressive neural network to frame-wise code distributions, trained on Kullback–Leibler divergence against consensus targets built by phoneme-aligned composition of real EMA segments. On the BCIC2020-IV benchmark (15 participants, five commands, 5250 trials) under five-fold cross-validation, the proposed mapper reached a held-out KL of 5.07, significantly below a soft-lookup baseline (9.54) and a hard-lookup baseline (12.26), with bootstrap intervals excluding zero. An oracle given the ground-truth command set a class-identity upper bound of 0.11. Per-command performance was balanced, and falsification controls collapsed under label shuffling and noise. On correctly identified trials the mapping approached this ceiling; the residual gap reflected the upstream EEG class signal rather than the articulatory framework, which we identify as the target for subject-invariant EEG representations.
| Original language | English |
|---|---|
| Article number | 101073 |
| Number of pages | 16 |
| Journal | Array |
| Volume | 31 |
| Early online date | 10 Jul 2026 |
| DOIs | |
| Publication status | E-pub ahead of print - 10 Jul 2026 |
Fingerprint
Dive into the research topics of 'A discrete articulatory codebook for imagined-speech EEG: Weakly supervised mapping to surrogate kinematics'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver