Artificial neural networks (ANNs) mimic the mechanisms of early visual processing in the human retina, but they do not emulate the complex mechanisms of facial expression interpretation in the human brain to the full extent.
In 1980, a Japanese computer scientist, Kunihiko Fukushima, invented the Neocognitron, the first deep convolutional neural network (CNN) (Fukushima, 2013). CNNs are a type of ANN that employs teaching signals for supervised learning (Cunningham, Cord, and Delany, 2008). The Neocognitron draws inspiration from the hierarchical properties of early visual processing in the human optic pathway, which extends from the retina to the primary visual cortex (Gupta et al., 2025). The CNN processes Facial Expression Recognition (FER) in a hierarchical multi-layered network, where ‘neurons within any given layer will only connect to a small region of the layer preceding it’ (O’Shea and Nash, 2015, p. 4).
Early visual processing in both CNNs and humans involves ‘parallel functional pathways’ (Sibille et al., 2022, p. 1). Reception of the stimulus is carried out analogously by the input layer in CNNs and the retina in the eye; information is then transferred respectively to early convolutional and pooling layers and to the lateral geniculate nucleus and visual cortex, both of which extract basic visual primitives (Derrington, 2001).
While the early visual encoding stage of FER is largely common in both CNNs and human networks, they follow different pathways during the more sophisticated interpretation of facial expressions (Zhao et al., 2024).
Firstly, human interpretation of facial expressions is very context specific, as the responsible brain area, our “social brain” (Hadders-Algra, 2022, p. 307), is influenced by unconscious biases like mood, gender and background colour.
Additionally, in the human brain, facial expressions are processed within the ventral face network. Not only does the network display redundant connections, which “do not follow a strict hierarchical organization” (Grill-Spector et al., 2017, p. 177), but it also interacts with dorsal stream regions and areas related to memory (Vuilleumier and Pourtois, 2007). Therefore, facial expression processing in the human brain involves complex interactions with both context and prior knowledge of faces.
In contrast, CNNs are trained through artificially developed teaching signals (Cunningham, Cord, and Delany, 2008), which are not able to incorporate the context specific details that unconsciously influence human perception of facial expressions. Moreover, CNNs process FER by hierarchically quantifying the changes in facial muscle movement from a neutral expression (Kim et al., 2019). Therefore, they do not accurately model the “highly dynamic and context sensitive” (Hadders-Algra, 2022, p. 311) mechanisms of human neural networks.
In conclusion, while the hierarchical early processing of visual stimuli in the retina can be emulated in CNNs, the higher level processing involved in evaluation of facial expressions entails complex interactions between areas of the social brain, beyond algorithmic modelling. It is arguably impossible for us humans to replicate our own ability to sense excitement in a fleeting smile or notice the loving sparkle in someone’s eyes — a paradox that reveals the miracle of conscious experience.
Bibliography
Bayet, L. and Nelson, C. A. (2020) ‘The neural architecture and developmental course of face processing’, in Neural Circuit and Cognitive Development. Academic Press, pp. 435–465.
Csikor, F. et al. (2025) ‘Top-down perceptual inference shaping the activity of early Visual Cortex’, Nature Communications, 16(1), p. 9998.
Cunningham, P., Cord, M. and Delany, S. J. (2008) ‘Supervised learning’, in M. Cord and P. Cunningham (eds) Machine Learning Techniques for Multimedia: Case Studies on Organization and Retrieval. Berlin, Heidelberg: Springer, pp. 21–49.
Dedalo, S. (2025) ‘Dare “corpo” all’intelligenza artificiale’, SapereScienza, (1), pp. 10–15.
Derrington, A. (2001) ‘The lateral geniculate nucleus’, Current Biology: CB, 11(16), pp. R635–637.
Flynn, C. et al. (2017) ‘Computational modeling of the passive and active components of the face’, in Biomechanics of Living Organs. Academic Press, pp. 377–394.
Fukushima, K. (2013) ‘Artificial vision by multi-layered neural networks: Neocognitron and its advances’, Neural Networks, 37, pp. 103–119.
Grill-Spector, K. et al. (2017) ‘The functional neuroanatomy of human face perception’, Annual Review of Vision Science, 3(1), pp. 167–196.
Gupta, M., Ireland, A. C. and Bordoni, B. (2025) ‘Neuroanatomy, Visual pathway’, in StatPearls. Treasure Island (FL): StatPearls Publishing. (Accessed: 15 December 2025.)
Hadders-Algra, M. (2022) ‘Human face and gaze perception is highly context specific and involves bottom-up and top-down neural processing’, Neuroscience and Biobehavioral Reviews, 132, pp. 304–323.
Kim, J.-H. et al. (2019) ‘Efficient facial expression recognition algorithm based on hierarchical deep neural network structure’, IEEE Access, 7, pp. 41273–41285.
O’Shea, K. and Nash, R. (2015) ‘An introduction to convolutional neural networks’. arXiv.
Sacks, O. W. (2010) ‘The Mind’s Eye’, in. New York: Alfred A. Knopf, p. 111.
Sibille, J. et al. (2022) ‘High-density electrode recordings reveal strong and specific connections between retinal ganglion cells and midbrain neurons’, Nature Communications, 13(1), p. 5218.
Vuilleumier, P. and Pourtois, G. (2007) ‘Distributed and interactive brain mechanisms during emotion face perception: Evidence from functional neuroimaging’, Neuropsychologia, 45(1), pp. 174–194.
Zhao, Z., Li, Y., Yang, J. and Ma, Y. (2024) ‘A lightweight facial expression recognition model for automated engagement detection’, Signal, Image and Video Processing, 18(4), pp. 3553–3563.
Author
