International Journal of Innovative Research in Computer and Communication Engineering

ISSN Approved Journal | Impact factor: 8.771 | ESTD: 2013 | Follows UGC CARE Journal Norms and Guidelines

| Monthly, Peer-Reviewed, Refereed, Scholarly, Multidisciplinary and Open Access Journal | High Impact Factor 8.771 (Calculated by Google Scholar and Semantic Scholar | AI-Powered Research Tool | Indexing in all Major Database & Metadata, Citation Generator | Digital Object Identifier (DOI) |


TITLE COGNITIVECANVAS: Reconstructing Visual Stimuli from Electroencephalogram (EEG) Signals Using Latent Diffusion Models
ABSTRACT Reconstructing visual experiences from non-invasive brain signals is a fundamental goal in Brain Computer Interface (BCI) technology. This project aims to create a system that decodes raw Electroencephalography (EEG) signals and reconstructs a photorealistic approximation of the image a person is viewing. This "visual mind-reading" moves beyond simple classification to the direct generation of pixel-level content. This task is exceptionally challenging due to the high noise and low spatial resolution inherent in EEG data. Current brain decoding research has had more success with fMRI, which offers high spatial resolution, but it is non-portable and expensive. EEG-based reconstruction is far more practical but also more difficult, representing an "extreme modality gap." Early attempts, often using Generative Adversarial Networks (GANs), have struggled to bridge this gap, producing low-resolution, blurry, or non-representative images that fail to capture the rich semantic detail of the original visual stimuli. This project introduces Cognitive Canvas, a novel dual-stage framework. First, a Transformer based EEG Encoder, trained with contrastive learning, will be designed to extract robust semantic feature vectors from the noisy, multi-channel time-series data. Second, a state-of-the-art Latent Diffusion Model (LDM) will be conditioned on these EEG feature vectors instead of text prompts to generate the final photorealistic image. This disentangled approach is designed to effectively translate abstract brain activity into a coherent visual space. The core implementation will use Python, PyTorch, and the Hugging Face diffusers library. The EEG Encoder will be a custom Transformer. Signal processing and feature extraction. Model alignment will enhance CLIP style contrastive learning. The system will be trained and evaluated on public EEG-visual datasets, primarily Mind Bigdata.
AUTHOR SYED ABDUL QUADEER KASHIF, DR. M. NAGARATNA Post-Graduate Student, Department of Computer Science Engineering, Data Science, Jawaharlal Nehru Technological University, Hyderabad, India Professor, Department of Computer Science Engineering, Jawaharlal Nehru Technological University, Hyderabad, India
VOLUME 185
DOI DOI: 10.15680/IJIRCCE.2026.1406057
PDF pdf/57_COGNITIVECANVAS Reconstructing Visual Stimuli from Electroencephalogram (EEG) Signals Using Latent Diffusion Models.pdf
KEYWORDS
References [1] D. Vivancos, "MindBigData 2022: A large dataset of brain signals," arXiv preprint arXiv:2301.01234, 2023.
[2] N. Kumari, S. Anwar, and V. Bhattacharjee, "Convolutional neural network-based visually evoked EEG classification model on MindBigData," International Journal of Cognitive Computing, vol. 12, pp. 45-56, 2021.
[3] R. Mishra, K. Sharma, and A. Bhavsar, "Visual brain decoding for short duration EEG signals," IEEE Transactions on Cognitive and Developmental Systems, vol. 14, pp. 1102-1114, 2021.
[4] Y. Takagi and S. Nishimoto, "High-resolution image reconstruction from human brain activity using latent diffusion models," in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2023, pp. 14453-14463.
[5] Anonymous, "Decoding EEG signals of visual brain representations with a CLIP based knowledge distillation," in International Conference on Learning Representations (ICLR) Conference Proceedings, 2024, pp. 23553-23565.
[6] R. Xiao et al., "Autoregressive visual decoding from EEG signals," arXiv preprint arXiv:2602.22555, 2026.
[7] J. Li et al., "Visual decoding and reconstruction via EEG embeddings with guided diffusion," Journal of Neural Engineering, vol. 22, pp. 1024-1035, 2025.
[8] C. Wang et al., "Reconstructing visual stimulus representation from EEG signals based on deep visual representation model," IEEE Transactions on Human-Machine Systems, vol. 54, no. 6, pp. 789-801, 2024.
[9] Z. Chen, J. Qing, T. Xiang, W. L. Yue, and J. H. Zhou, "Seeing beyond the brain: Conditional diffusion model with sparse masked Modelling for vision decoding," in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2023, pp. 22710-22720.
[10] R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer, "High-resolution image synthesis with latent diffusion models," in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022, pp. 10684-10695.
[11] A. Radford et al., "Learning transferable visual models from natural language supervision," in International Conference on Machine Learning (ICML), 2021, pp. 8748-8763.
[12] A. Vaswani et al., "Attention is all you need," Advances in Neural Information Processing Systems (NeurIPS), vol. 30, 2017.
[13] M. Heusel, H. Ramsauer, T. Unterthiner, B. Nessler, and S. Hochreiter, "GANs trained by a two time-scale update rule converge to a local Nash equilibrium," Advances in Neural Information Processing Systems (NeurIPS), vol. 30, 2017.
image
Copyright © IJIRCCE 2020.All right reserved