Abstract
Affective computing has grown into a field that spans emotion theory, multimodal data, computational modeling, human–AI interaction, social applications, and ethical evaluation, yet most existing surveys address only one or two of these layers at a time and fail to capture how the layers depend on one another. Rather than offering another comprehensive survey, this position paper reorganizes the field around a six-layer pipeline and reads the literature as a chain of dependencies in which upstream design choices propagate into downstream evaluation, deployment, and governance. Our central diagnosis is that four silo-bridge patterns recur across the field: disconnections at the theory/data (Pattern 1), modeling/interaction (Pattern 2), technology/ethics (Pattern 3), and data/social-application (Pattern 4) boundaries. These are not incidental implementation problems but recurring failure modes rooted in the pipeline structure itself, forming a cascade in which upstream patterns structurally enable downstream ones. We argue that this dependency chain is not merely a technical liability but a socio-philosophical one: as upstream inconsistencies propagate downstream, they threaten the individual’s authority to interpret their own emotions (emotional agency), pull affective expression toward machine-legible norms (algorithmic conformity), and relocate the power to define and act on emotion from persons to the institutions that operate these systems (institutional power). As a prescriptive program, we present five integrated design criteria: theory-explicit affective representation (DC1), affective reasoning with bounded intervention (DC2), longitudinal interaction-in-the-loop evaluation (DC3), deployment-specific affective accountability (DC4), and user-retained interpretive authority (DC5). The five criteria are designed as an interlocking response to the cascade rather than as independent recommendations. We situate affective AI not as a mere recognition or generation technology but as a sociotechnical system that can intervene in human emotional interpretation, interaction, and social decision-making, and call upon future research to move toward systems that are more coherent, more responsible, and more respectful of human emotional agency.
Similar content being viewed by others
Data availability
No datasets were generated or analyzed during the current study.
Materials availability
Not applicable.
Code availability
Not applicable.
References
Ahir A, Gohokar V (2019) Driver inattention monitoring system: a review. In: Proceedings of the 2019 international conference on innovative trends and advances in engineering and technology (ICITAET), pp 188–194. https://doi.org/10.1109/ICITAET47105.2019.9170249
Andalibi N, Ingber AS (2025) Public perceptions about emotion ai use across contexts in the united states. In: Proceedings of the 2025 CHI conference on human factors in computing systems, pp 1–16. https://doi.org/10.1145/3706598.3713501
Averill FR (1980) A constructivist view of emotion. In: Theories of emotion, pp 305–339. https://doi.org/10.1016/B978-0-12-558701-3.50018-1
Barker D, Tippireddy MKR, Farhan A, Ahmed B (2025) Ethical considerations in emotion recognition research. Psychol Int 7(2):43. https://doi.org/10.3390/psycholint7020043
Barrett LF (2006) Are emotions natural kinds? Perspect Psychol Sci 1(1):28–58. https://doi.org/10.1111/j.1745-6916.2006.00003.x
Barrett LF (2017) The theory of constructed emotion: an active inference account of interoception and categorization. Soc Cognit Affect Neurosci 12(1):1–23. https://doi.org/10.1093/scan/nsw154
Barrett LF, Adolphs R, Marsella S, Martinez AM, Pollak SD (2019) Emotional expressions reconsidered: challenges to inferring emotion from human facial movements. Psychol Sci Public Interest 20(1):1–68. https://doi.org/10.1177/1529100619832930
Bender EM, Friedman B (2018) Data statements for natural language processing: toward mitigating system bias and enabling better science. Trans Assoc Comput Linguist 6:587–604. https://doi.org/10.1162/tacl_a_00041
Benjamin R (2019) Race after technology: abolitionist tools for the new jim code. Polity
Bickmore TW, Picard RW (2005) Establishing and maintaining long-term human-computer relationships. ACM Trans Comput Hum Interact 12(2):293–327. https://doi.org/10.1145/1067860.1067867
Bohus D, Horvitz E (2009) Models for multiparty engagement in open-world dialog. In: Proceedings of SIGDIAL 2009, pp 225–234
Bostan LAM, Klinger R (2018) An analysis of annotated corpora for emotion classification in text. In: Proceedings of the 27th international conference on computational linguistics, pp 2104–2119
Breazeal C (2003) Emotion and sociable humanoid robots. Int J Hum Comput Stud 59(1–2):119–155. https://doi.org/10.1016/S1071-5819(03)00018-1
Breazeal C (2004) Designing sociable robots. MIT Press. https://doi.org/10.7551/mitpress/2376.001.0001
Broekens J, Hilpert B, Verberne S, Baraka K, Gebhard P, Plaat A (2023) Fine-grained affective processing capabilities emerging from large language models. In: Proceedings of the international conference on affective computing and intelligent interaction, pp 1–8. https://doi.org/10.1109/ACII59096.2023.10388177
Buolamwini J, Gebru T (2018) Gender shades: intersectional accuracy disparities in commercial gender classification. In: Proceedings of the 1st conference on fairness, accountability and transparency. Proceedings of machine learning research, vol 81, pp 77–91
Burkhardt F, Paeschke A, Rolfes M, Sendlmeier WF, Weiss B (2005) A database of German emotional speech. In: Proceedings of the annual conference of the international speech communication association, vol 5. ISCA, Lisbon, pp 1517–1520. https://doi.org/10.21437/Interspeech.2005-446
Busso C, Bulut M, Lee C-C, Kazemzadeh A, Mower E, Kim S, Chang JN, Lee S, Narayanan SS (2008) IEMOCAP: interactive emotional dyadic motion capture database. Lang Resour Eval 42:335–359. https://doi.org/10.1007/S10579-008-9076-6
Calvo RA, D’Mello S (2010) Affect detection: an interdisciplinary review of models, methods, and their applications. IEEE Trans Affect Comput. https://doi.org/10.1109/T-AFFC.2010.1
Cambria E (2016) Affective computing and sentiment analysis. IEEE Intell Syst 31(2):102–107. https://doi.org/10.1109/MIS.2016.31
Cao H, Cooper DG, Keutmann MK, Gur RC, Nenkova A, Verma R (2014) CREMA-D: crowd-sourced emotional multimodal actors dataset. IEEE Trans Affect Comput 5(4):377–390. https://doi.org/10.1109/TAFFC.2014.2336244
Cassell J, Sullivan J, Prevost S, Churchill EF (2000) Embodied conversational agents. MIT Press. https://doi.org/10.7551/mitpress/2697.001.0001
Chalmers DJ (1995) Facing up to the problem of consciousness. J Conscious Stud 2(3):200–219. https://doi.org/10.1093/acprof:oso/9780195311105.003.0001
Chandra NA, Murtfeldt R, Qiu L, Karmakar A, Lee H, Tanumihardja E, Farhat K, Caffee B, Paik S, Lee C, Choi J, Kim A, Etzioni O (2025) Deepfake-Eval-2024: a multi-modal in-the-wild benchmark of deepfakes circulated in 2024. https://doi.org/10.48550/arXiv.2503.02857 arXiv: 2503.02857
Chavan V, Cenaj A, Shen S, Bar A, Binwani S, Del Becaro T, Funk M, Greschner L, Hung R, Klein S, Kleiner R, Krause S, Olbrych S, Parmar V, Sarafraz J, Soroko D, Withanage Don D, Zhou C, Vu HTD, Semnani P, Weinhardt D, Andre E, Kr ger J, Fresquet X (2025) Feeling machines: ethics, culture, and the rise of emotional ai. https://doi.org/10.48550/arXiv.2506.12437 arXiv:2506.12437
Cheng Z, Cheng Z-Q, He J-Y, Sun J, Wang K, Lin Y, Lian Z, Peng X, Hauptmann AG (2024) Emotion-LLaMA: multimodal emotion recognition and reasoning with instruction tuning. In: Proceedings of the 38th international conference on neural information processing systems, pp 110805–110853. https://doi.org/10.5555/3737916.3741434
Damasio A (1994) Descartes’ error: emotion, reason, and the human brain. Putnam
Davidson RJ (2004) What does the prefrontal cortex“do’’in affect: perspectives on frontal EEG asymmetry research. Biol Psychol 67(1–2):219–234. https://doi.org/10.1016/j.biopsycho.2004.03.008
Davis MH (1983) Measuring individual differences in empathy: evidence for a multidimensional approach. J Pers Soc Psychol 44(1):113–126. https://doi.org/10.1037/0022-3514.44.1.113
Davison AK, Lansley C, Costen N, Tan K, Yap MH (2018) Samm: a spontaneous micro-facial movement dataset. IEEE Trans Affect Comput 9(1):116–129. https://doi.org/10.1109/TAFFC.2016.2573832
De Choudhury M, Gamon M, Counts S, Horvitz E (2013) Predicting depression via social media. In: Proceedings of the international AAAI conference on web and social media, vol 7, pp 128–137. https://doi.org/10.1609/icwsm.v7i1.14432
Decety J, Jackson PL (2004) The functional architecture of human empathy. Behav Cogn Neurosci Rev 3(2):71–100. https://doi.org/10.1177/1534582304267
Défossez A, Mazaré L, Orsini M, Royer A, Pérez P, Jégou H, Grave E, Zeghidour N (2024) Moshi: a speech-text foundation model for real-time dialogue. https://doi.org/10.48550/arXiv.2410.00037 arXiv:2410.00037
Demszky D, Movshovitz-Attias D, Ko J, Cowen A, Nemade G, Ravi S (2020) GoEmotions: a dataset of fine-grained emotions. In: Proceedings of the 58th annual meeting of the association for computational linguistics, pp 4040–4054. https://doi.org/10.18653/v1/2020.acl-main.372
DeVault D, Artstein R, Benn G, Dey T, Fast E, Gainer A, Georgila K, Gratch J, Hartholt A, Lhommet M, Lucas G, Marsella S, Morbini F, Nazarian A, Scherer S, Stratou G, Suri A, Traum D, Wood R, Xu Y, Rizzo A, Morency L-P (2014) Simsensei kiosk: a virtual human interviewer for healthcare decision support. In: Proceedings of the 2014 international conference on autonomous agents and multi-agent systems, pp 1061–1068. https://doi.org/10.5555/2615731.2617415
Devlin J, Chang M-W, Lee K, Toutanova K (2019) BERT: pre-training of deep bidirectional transformers for language understanding. In: Proceedings of the 2019 conference of the North American chapter of the association for computational linguistics: human language technologies, pp 4171–4186. https://doi.org/10.18653/v1/N19-1423
Dinan E, Roller S, Shuster K, Fan A, Auli M, Weston J (2019) Wizard of wikipedia: knowledge-powered conversational agents. In: Proceedings of the international conference on learning representations. https://doi.org/10.48550/arXiv.1811.01241
D’Mello S (2013) A selective meta-analysis on the relative incidence of discrete affective states during learning with technology. J Educ Psychol 105(4):1082–1099. https://doi.org/10.1037/a0032674
D’Mello SK, Graesser A (2010) Multimodal semi-automated affect detection from conversational cues, gross body language, and facial features. User Model User-Adap Interact 20:147–187. https://doi.org/10.1007/s11257-010-9074-4
Duffy BR (2003) Anthropomorphism and the social robot. Robot Auton Syst 42(3–4):177–190. https://doi.org/10.1016/S0921-8890(02)00374-3
Ekman P (1992) An argument for basic emotions. Cognit Emot 6(3–4):169–200. https://doi.org/10.1080/02699939208411068
Ekman P, Friesen WV (1978) Facial action coding system: a technique for the measurement of facial movement. Consulting Psychologists Press. https://doi.org/10.1037/t27734-000
Elfenbein HA, Ambady N (2002) On the universality and cultural specificity of emotion recognition: a meta-analysis. Psychol Bull 128(2):203–235. https://doi.org/10.1037/0033-2909.128.2.203
Emmelkamp PMG, Meyerbrker K (2021) Virtual reality therapy in mental health. Annu Rev Clin Psychol 17:495–519. https://doi.org/10.1146/annurev-clinpsy-081219-115923
Fang CM, Liu AR, Danry V, Lee E, Chan SWT, Pataranutaporn P, Maes P, Phang J, Lampe M, Ahmad L, Agarwal S (2024) How AI and human behaviors shape psychosocial effects of chatbot use: a longitudinal randomized controlled study. https://doi.org/10.48550/arXiv.2503.17473 arXiv: 2503.17473
Fischer T, Biemann C (2024) Exploring large language models for qualitative data analysis. In: Proceedings of the 4th international conference on natural language processing for digital humanities, pp 423–437. https://aclanthology.org/2024.nlp4dh-1.41/
Fitzpatrick KK, Darcy A, Vierhile M (2017) Delivering cognitive behavior therapy to young adults with symptoms of depression and anxiety using a fully automated conversational agent (Woebot): a randomized controlled trial. JMIR Mental Health 4(2):19. https://doi.org/10.2196/mental.7785
Foucault M (1977) Discipline and punish: the birth of the prison. Vintage Books
Gabriel S, Puri I, Xu X, Malgaroli M, Ghassemi M (2024) Can AI relate: testing large language model response for mental health support. In: Findings of the association for computational linguistics: EMNLP 2024, pp 2206–2221. https://doi.org/10.18653/v1/2024.findings-emnlp.120
Gebru T, Morgenstern J, Vecchione B, Vaughan JW, Wallach H, Daum H III, Crawford K (2021) Datasheets for datasets. Commun ACM 64(12):86–92
Geirhos R, Jacobsen J-H, Michaelis C, Zemel R, Brendel W, Bethge M, Wichmann FA (2020) Shortcut learning in deep neural networks. Nat Mach Intell 2:665–673. https://doi.org/10.1038/s42256-020-00257-z
Gergen KJ (1985) The social constructionist movement in modern psychology. Am Psychol 40(3):266–275. https://doi.org/10.1037/0003-066X.40.3.266
Ghosal D, Majumder N, Poria S, Chhaya N, Gelbukh A (2019) DialogueGCN: a graph convolutional neural network for emotion recognition in conversation. In: Proceedings of the 2019 conference on empirical methods in natural language processing and the 9th international joint conference on natural language processing, pp 154–164. https://doi.org/10.18653/v1/D19-1015
Gilardi F, Alizadeh M, Kubli M (2023) ChatGPT outperforms crowd workers for text-annotation tasks. Proc Natl Acad Sci. https://doi.org/10.1073/pnas.2305016120
Goodfellow IJ, Erhan D, Carrier PL, Courville A, Mirza M, Hamner B, Cukierski W, Tang Y, Thaler D, Lee D-H, Zhou Y, Ramaiah C, Feng F, Li R, Wang X, Athanasakis D, Shawe-Taylor J, Milakov M, Park J, Ionescu R, Popescu M, Grozea C, Bergstra J, Xie J, Romaszko L, Xu B, Chuang Z, Bengio Y (2013) Challenges in representation learning: a report on three machine learning contests. In: Proceedings of the international conference on neural information processing, pp 117–124. https://doi.org/10.1007/978-3-642-42051-1_16
Gross JJ (1998) The emerging field of emotion regulation: an integrative review. Rev Gen Psychol 2(3):271–299. https://doi.org/10.1037/1089-2680.2.3.271
Gross JJ (2015) Emotion regulation: current status and future prospects. Psychol Inq 26(1):1–26. https://doi.org/10.1080/1047840X.2014.940781
Gross JJ, Levenson RW (1995) Emotion elicitation using films. Cognit Emot 9(1):87–108. https://doi.org/10.1080/02699939508408966
Guo Z, Lai A, Thygesen JH, Farrington J, Keen T, Li K (2024) Large language models for mental health applications: systematic review. JMIR Mental Health 11:57400. https://doi.org/10.2196/57400
Hagendorff T (2020) The ethics of AI ethics: an evaluation of guidelines. Mind Mach 30(1):99–120. https://doi.org/10.1007/s11023-020-09517-8
Halkiopoulos C, Gkintoni E, Aroutzidis A, Antonopoulou H (2025) Advances in neuroimaging and deep learning for emotion detection: a systematic review of cognitive neuroscience and algorithmic innovations. Diagnostics 15(4):456. https://doi.org/10.3390/diagnostics15040456
He H, Garcia EA (2009) Learning from imbalanced data. IEEE Trans Knowl Data Eng 21(9):1263–1284. https://doi.org/10.1109/TKDE.2008.239
Hegde K, Jayalath H (2025) Emotions in the loop: a survey of affective computing for emotional support. https://doi.org/10.48550/arXiv.2505.01542 arXiv:2505.01542
Hochreiter S, Schmidhuber J (1997) Long short-term memory. Neural Comput 9(8):1735–1780. https://doi.org/10.1162/neco.1997.9.8.1735
Hochschild AR (1979) The managed heart: commercialization of human feeling. University of California Press
Huang C-ZA, Vaswani A, Uszkoreit J, Shazeer N, Simon I, Hawthorne C, Dai AM, Hoffman MD, Dinculescu M, Eck D (2018) Music transformer. https://doi.org/10.48550/arXiv.1809.04281 arXiv:1809.04281
Huang X, Hong X, Mao Q, Zheng W, Dhall A (2024) A survey on deep learning for group-level emotion recognition. IEEE Trans Comput Soc Syst 13(2):2475–2500. https://doi.org/10.1109/TCSS.2025.3638859
Hutto C, Gilbert E (2014) VADER: a parsimonious rule-based model for sentiment analysis of social media text. In: Proceedings of the international AAAI conference on web and social media, vol 8, pp 216–225
Imel ZE, Caperton DD, Tanana M, Atkins DC (2017) Technology-enhanced human interaction in psychotherapy. J Couns Psychol 64(4):385–393. https://doi.org/10.1037/cou0000213
Indrasiri PL, Kashyap B, Kolambahewage C, Nakisa B, Ijaz K, Pathirana PN (2024) VR based emotion recognition using deep multimodal fusion with biosignals across multiple anatomical domains. https://doi.org/10.48550/arXiv.2412.02283 arXiv:2412.02283
Inkster B, Sarda S, Subramanian V (2018) Delivering cognitive behavior therapy to young adults with symptoms of depression and anxiety using a fully automated conversational agent (Woebot): a randomized controlled trial. JMIR Mental Health 4(2):19. https://doi.org/10.2196/12106
Inoshita K (2024) Sentiment analysis of Japanese twitter users regarding the Ukraine-Russia War and its implications for security policy. In: 2024 11th international conference on information technology, computer, and electrical engineering, pp 338–343. https://doi.org/10.1109/ICITACEE62763.2024.10762783
Inoshita K, Harada R (2026) PersonaGen: persona-based synthetic data generation using multi-stage conditioning with large language models for emotion recognition. Int J Act Behav Comput 1:1–18. https://doi.org/10.60401/ijabc.133
Inoshita K, Mizuno S (2026) World model inspired sarcasm reasoning with large language model agents. https://doi.org/10.48550/arXiv.2512.24329
Inoshita K, Tomisu H, Zhou X, Kawai A, Yada K (2026) KDDA: a knowledge-driven domain and diversity alignment framework for emotion data generation with large language models. Int J Act Behav Comput 1:1–24
Inoshita K, Zhou X, Kawai A, Yada K (2026b) LLMs capture emotion labels, not emotion uncertainty: distributional analysis and calibration of human-LLM judgment gaps. https://doi.org/10.48550/arXiv.2604.27345
Irfan B, Kuoppamäki S, Skantze G (2024) Recommendations for designing conversational companion robots with older adults through foundation models. Front Robot AI. https://doi.org/10.3389/frobt.2024.1363713
Jiang W, Windl M, Tag B, Sarsenbayeva Z, Mayer S (2024) An immersive and interactive vr dataset to elicit emotions. IEEE Trans Visual Comput Graph 30(11):7343–7353. https://doi.org/10.1109/TVCG.2024.3456202
Jin Y, Liu J, Li P, Wang B, Yan Y, Zhang H, Ni C, Wang J, Li Y, Bu Y, Wang Y (2025) The applications of large language models in mental health: scoping review. J Med Internet Res 27:69284. https://doi.org/10.2196/69284
Jobin A, Ienca M, Vayena E (2019) The global landscape of AI ethics guidelines. Nat Mach Intell 1:389–399. https://doi.org/10.1038/s42256-019-0088-2
Johnson KT, Narain J, Thomas Q, Maes P, Picard RW (2023) Recanvo: a database of real-world communicative and affective nonverbal vocalizations. Sci Data. https://doi.org/10.1038/s41597-023-02405-7
Ju Z, Wang Y, Shen K, Tan X, Xin D, Yang D, Liu Y, Leng Y, Song K, Tang S, Wu Z, Qin T, Li X, Ye W, Zhang S, Bian J, He L, Li J, Zhao S (2024) NaturalSpeech 3: zero-shot speech synthesis with factorized codec and diffusion models. In: Proceedings of the 41st international conference on machine learning, pp 22605–22623. https://doi.org/10.5555/3692070.3692979
Kaplan AD, Kessler TT, Brill JC, Hancock PA (2021) Trust in artificial intelligence: meta-analytic findings. J Hum Factors Ergon Soc. https://doi.org/10.1177/00187208211013988
Khalil HA, Hammad SA, Munim HEAE, Maged SA (2023) Low-cost driver monitoring system using deep learning. IEEE Access 13:14151–14164. https://doi.org/10.1109/ACCESS.2025.3530296
Khan UA, Xu Q, Liu Y, Lagstedt A, Alamäki A, Kauttonen J (2024) Exploring contactless techniques in multimodal emotion recognition: insights into diverse applications, challenges, solutions, and prospects. Multimed Syst. https://doi.org/10.1007/s00530-024-01302-2
Kim Y (2014) Convolutional neural networks for sentence classification. In: Proceedings of the 2014 conference on empirical methods in natural language processing, pp 1746–1751. https://doi.org/10.3115/v1/D14-1181
Kim RS (2026) Formal and computational foundations for implementing affective sovereignty in emotion AI systems. Discov Artif Intell 6(235):501–507. https://doi.org/10.1007/s44163-026-01000-0
Koelstra S, Mühl C, Soleymani M, Lee J-S, Yazdani A, Ebrahimi T, Pun T, Nijholt A, Patras I (2012) DEAP: a database for emotion analysis using physiological signals. IEEE Trans Affect Comput 3(1):18–31. https://doi.org/10.1109/T-AFFC.2011.15
Kumar MJD, Rao MS, Narendra KC (2025) Multimodal emotion recognition: a comprehensive survey of datasets, methods, and applications. IEEE Access 13:201067–201097. https://doi.org/10.1109/ACCESS.2025.3636186
Lang PJ, Bradley MM, Cuthbert BN (2005) International affective picture system (iaps): affective ratings of pictures and instruction manual. Technical Report Technical Report A-6. University of Florida, Gainesville
Lazarus RS (1991) Emotion and adaptation. Oxford University Press
LeDoux JE (1998) The emotional brain: the mysterious underpinnings of emotional life. Weidenfeld & Nicolson
Lei S, Dong G, Wang X, Wang K, Qiao R, Wang S (2024a) InstructERC: reforming emotion recognition in conversation with multi-task retrieval-augmented large language models. arXiv:2309.11911. https://doi.org/10.48550/arXiv.2309.11911
Lei S, Zhou Y, Tang B, Lam MWY, Liu F, Liu H, Wu J, Kang S, Wu Z, Meng H (2024b) SongCreator: lyrics-based Universal Song Generation. In: Proceedings of the 38th international conference on neural information processing systems, pp 80107–80140. https://doi.org/10.52202/079017-2546
Leite I, Castellano G, Pereira A, Martinho C, Paiva A (2014) Empathic robots for long-term interaction. Int J Soc Robot 6:329–341. https://doi.org/10.1007/s12369-014-0227-1
Li Y, Su H, Shen X, Li W, Cao Z, Niu S (2017) DailyDialog: a manually labelled multi-turn dialogue dataset. In: Proceedings of the 8th international joint conference on natural language processing, pp 986–995
Lian Z, Sun H, Chen L, Sun H, Sun L, Ren Y, Cheng Z, Liu B, Liu R, Peng X, Yi J, Tao J (2025) AffectGPT: a new dataset, model, and benchmark for emotion understanding with multimodal large language models. In: Proceedings of the 42nd international conference on machine learning, pp 36993–37014
Lin Z, Madotto A, Shin J, Xu P, Fung P (2019) MoEL: mixture of empathetic listeners. In: Proceedings of the 2019 conference on empirical methods in natural language processing and the 9th international joint conference on natural language processing, pp 121–132. https://doi.org/10.18653/v1/D19-1012
Lin H, Czarnek G, Lewis B, White JP, Berinsky AJ, Costello T, Pennycook G, Rand DG (2025) Persuading voters using human-artificial intelligence dialogues. Nature 648:394–401. https://doi.org/10.1038/s41586-025-09771-9
Liu Z, Shen Y, Lakshminarasimhan VB, Liang PP, Zadeh AB, Morency L-P (2018) Efficient low-rank multimodal fusion with modality-specific factors. In: Proceedings of the 56th annual meeting of the association for computational linguistics, pp 2247–2256. https://doi.org/10.18653/v1/P18-1209
Liu S, Zheng C, Demasi O, Sabour S, Li Y, Yu Z, Jiang Y, Huang M (2021) Towards emotional support dialog systems. In: Proceedings of the 59th annual meeting of the association for computational linguistics and the 11th international joint conference on natural language processing, pp 3469–3483. https://doi.org/10.18653/v1/2021.acl-long.269
Liu Y, Dai W, Feng C, Wang W, Yin G, Zeng J, Shan S (2022) MAFW: a large-scale, multi-modal, compound affective database for dynamic facial expression recognition in the wild. In: Proceedings of the 30th ACM international conference on multimedia. Association for Computing Machinery, pp 24–32. https://doi.org/10.1145/3503161.3548190
Livingstone SR, Russo FA (2018) The Ryerson audio-visual database of emotional speech and song (ravdess): a dynamic, multimodal set of facial and vocal expressions in north American English. PLoS ONE 13(5):0196391. https://doi.org/10.1371/journal.pone.0196391
Lotfian R, Busso C (2019) Building naturalistic emotionally balanced speech corpus by retrieving emotional speech from existing podcast recordings. IEEE Trans Affect Comput 10(4):471–483. https://doi.org/10.1109/TAFFC.2017.2736999
Lubold N, Walker E, Pon-Barry H, Ogan A (2018) Automated pitch convergence improves learning in a social, teachable robot for middle school mathematics. In: Proceedings of the 19th international conference on artificial intelligence in education. Lecture notes in computer science, vol 10947, pp 282–296. https://doi.org/10.1007/978-3-319-93843-1_21
Lucey P, Cohn JF, Kanade T, Saragih J, Ambadar Z, Matthews I (2010) The extended Cohn-Kanade dataset (ck+): a complete dataset for action unit and emotion-specified expression. In: 2010 IEEE computer society conference on computer vision and pattern recognition workshops, pp 94–101. https://doi.org/10.1109/CVPRW.2010.5543262
Lutz CA (1988) Unnatural emotions: everyday sentiments on a micronesian atoll and their challenge to western theory. University of Chicago Press
Ma Z, Zheng Z, Ye J, Li J, Gao Z, Zhang S, Chen X (2024) emotion2vec: self-supervised pre-training for speech emotion representation. In: Findings of the association for computational linguistics: ACL 2024, pp 15747–15760. https://doi.org/10.18653/v1/2024.findings-acl.931
Ma F, Yuan Y, Xie Y, Ren H, Liu I, He Y, Ren F, Yu FR, Ni S (2025) Generative technology for human emotion recognition: a scoping review. Inf Fusion 115:102753. https://doi.org/10.1016/j.inffus.2024.102753
Maas AL, Daly RE, Pham PT, Huang D, Ng AY, Potts C (2011) Learning word vectors for sentiment analysis. In: Proceedings of the 49th annual meeting of the association for computational linguistics: human language technologies, pp 142–150
Madiega T (2024) Artificial intelligence act. Briefing PE 698.792. European Parliamentary Research Service
Maeda T, Quan-Haase A (2024) When human-AI interactions become parasocial: agency and anthropomorphism in affective design. In: Proceedings of the 2024 ACM conference on fairness, accountability, and transparency. https://doi.org/10.1145/3630106.3658956
Majumder N, Poria S, Hazarika D, Mihalcea R, Gelbukh A, Cambria E (2019) DialogueRNN: an attentive RNN for emotion detection in conversations. In: Proceedings of the thirty-third AAAI conference on artificial intelligence, pp 6818–6825. https://doi.org/10.1609/aaai.v33i01.33016818
Mancuso V, Borghesi F, Chirico A, Bruni F, Sarcinella ED, Pedroli E, Cipresso P (2024) Iavrs international affective virtual reality system: psychometric assessment of 360 images by using psychophysiological data. Sensors 24(13):4204. https://doi.org/10.3390/s24134204
Maples B, Cerit M, Vishwanath A, Pea R (2024) Loneliness and suicide mitigation for students using GPT3-enabled chatbots. NPJ Ment Health Res. https://doi.org/10.1038/s44184-023-00047-6
McColl D, Hong A, Hatakeyama N, Nejat G, Benhabib B (2016) A survey of autonomous human affect detection methods for social robots engaged in natural HRI. J Intell Robot Syst 82:101–133. https://doi.org/10.1007/s10846-015-0259-2
McKeown G, Valstar M, Cowie R, Pantic M, Schroder M (2012) The SEMAINE database: annotated multimodal records of emotionally colored conversations between a person and a limited agent. IEEE Trans Affect Comput 3(1):5–17. https://doi.org/10.1109/T-AFFC.2011.25
McStay A (2018) Emotional AI: the rise of empathic media. SAGE Publications. https://doi.org/10.4135/9781526451293
Mehrabian A, Russell JA (1974) An approach to environmental psychology. MIT Press
Miner AS, Milstein A, Schueller S, Hegde R, Mangurian C, Linos E (2016) Smartphone-based conversational agents and responses to questions about mental health, interpersonal violence, and physical health. JAMA Intern Med 176(5):619–625
Mitchell M, Wu S, Zaldivar A, Barnes P, Vasserman L, Hutchinson B, Spitzer E, Raji ID, Gebru T (2019) Model cards for model reporting. In: Proceedings of the conference on fairness, accountability, and transparency, pp 220–229. https://doi.org/10.1145/3287560.3287596
Mittelstadt B (2019) Principles alone cannot guarantee ethical AI. Nat Mach Intell 1:501–507. https://doi.org/10.1038/s42256-019-0114-4
Mohammad SM (2022) Ethics sheet for automatic emotion recognition and sentiment analysis. Comput Linguist 48(2):239–278. https://doi.org/10.1162/coli_a_00433
Mohammad SM, Turney PD (2012) Crowdsourcing a word-emotion association lexicon. Comput Intell. https://doi.org/10.1111/j.1467-8640.2012.00460.x
Mohamed S, Png M-T, Isaac W (2020) Decolonial AI: decolonial theory as sociotechnical foresight in artificial intelligence. Philos Technol 33:659–684. https://doi.org/10.1007/s13347-020-00405-8
Mollahosseini A, Hasani B, Mahoor MH (2019) AffectNet: a database for facial expression, valence, and arousal computing in the wild. IEEE Trans Affect Comput 10(1):18–31. https://doi.org/10.1109/TAFFC.2017.2740923
Moon A-S, Kim H, Park Y-C, Lee J (2026) A survey on multimodal emotion recognition: methods, datasets, and future directions. Comput Mater Continua 1:1. https://doi.org/10.32604/cmc.2026.076411
Moors A, Ellsworth PC, Scherer KR, Frijda NH (2013) Appraisal theories of emotion: state of the art and future development. Emot Rev 5(2):119–124. https://doi.org/10.1177/1754073912468165
Mori H, Nishino H (2025) End-to-end conversational speech synthesis with controllable emotions in the dimensions of pleasantness and arousal. Acoust Sci Technol 46(1):70–77. https://doi.org/10.1250/ast.e24.13
Munezero MD, Montero CS, Sutinen E, Pajunen J (2014) Are they different? Affect, feeling, emotion, sentiment, and opinion detection in text. IEEE Trans Affect Comput 5(2):101–111. https://doi.org/10.1109/TAFFC.2014.2317187
Nabulsi J (2025) Affective sovereignty: a decolonising politics of emotion in palestine. Rev Int Stud. https://doi.org/10.1017/S0260210525100880
Niedenthal PM (2007) Embodying emotion. Science 316(5827):1002–1005. https://doi.org/10.1126/science.1136930
Norcross JC (2011) Psychotherapy relationships that work: evidence-based responsiveness. In: Psychotherapy relationships that work: evidence-based responsiveness, 2nd edn. Oxford University Press
Northcutt CG, Athalye A, Mueller J (2021) Pervasive label errors in test sets destabilize machine learning benchmarks. In: Proceedings of the 35th international conference on neural information processing systems, track on datasets and benchmarks
Ocumpaugh J, Baker RS, Gowda SM, Heffernan NT, Heffernan C (2014) Population validity for educational data mining models: a case study in affect detection. Br J Edu Technol 45:487–501. https://doi.org/10.1111/bjet.12156
Omarov B, Narynov S, Zhumanov Z (2022) Artificial intelligence-enabled chatbots in mental health: a systematic review. Comput Mater Continua 74(3):5105–5122. https://doi.org/10.32604/cmc.2023.034655
Ong DC (2021) An ethical framework for guiding the development of affectively-aware artificial intelligence. In: Proceedings of the 9th international conference on affective computing and intelligent interaction, pp 1–8. https://doi.org/10.1109/ACII52823.2021.9597441
Oorloff T, Koppisetti S, Bonettini N, Solanki D, Colman B, Yacoob Y, Shahriyari A, Bharaj G (2024) AVFF: audio-visual feature fusion for video deepfake detection. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition 2024, pp 27102–27112
Ortony A, Clore GL, Collins A (1988) The cognitive structure of emotions. Cambridge University Press. https://doi.org/10.1017/CBO9780511571299
Pang B, Lee L (2008) Opinion mining and sentiment analysis. Found Trends Inf Retr 2(1–2):1–135. https://doi.org/10.1561/1500000011
Park CY, Cha N, Kang S, Kim A, Khandoker AH, Hadjileontiadis L, Oh A, Jeong Y, Lee U (2020) K-EmoCon, a multimodal sensor dataset for continuous emotion recognition in naturalistic conversations. Sci Data. https://doi.org/10.1038/s41597-020-00630-y
Parra-Gallego LF, Arias-Vergara T, Orozco-Arroyave JR (2025) Multimodal evaluation of customer satisfaction from voicemails using speech and language representations. Dig Signal Process 156(B):104820. https://doi.org/10.1016/j.dsp.2024.104820
Pei G, Haiying L, Lu Y, Wang Y, Hua S, Thihao L (2024) Affective computing: recent advances, challenges, and future trends. Intell Comput. https://doi.org/10.34133/icomputing.0076
Pelachaud C (2009) Studies on gesture expressivity for a virtual agent. Speech Commun 51(7):630–639. https://doi.org/10.1016/j.specom.2008.04.009
Picard RW (1997) Affective computing. MIT Press. https://doi.org/10.7551/mitpress/1140.001.0001
Picard RW, Vyzas E, Healey J (2001) Toward machine emotional intelligence: analysis of affective physiological state. IEEE Trans Pattern Anal Mach Intell 23(10):1175–1191. https://doi.org/10.1109/34.954607
Plutchik R (1980) A general psychoevolutionary theory of emotion. In: Plutchik R, Kellerman H (eds) Theories of emotion, vol 1. Academic Press, pp 3–33. https://doi.org/10.1016/B978-0-12-558701-3.50007-7
Poria S, Cambria E, Bajpai R, Hussain A (2017) A review of affective computing: from unimodal analysis to multimodal fusion. Inf Fusion 37:98–125. https://doi.org/10.1016/j.inffus.2017.02.003
Poria S, Majumder N, Mihalcea R, Hovy E (2019) Emotion recognition in conversation: research challenges, datasets, and recent advances. IEEE Access 7:100943–100953. https://doi.org/10.1109/ACCESS.2019.2929050
Poria S, Hazarika D, Majumder N, Naik G, Cambria E, Mihalcea R (2019b) MELD: a multimodal multi-party dataset for emotion recognition in conversations. In: Proceedings of the 57th annual meeting of the association for computational linguistics, pp 527–536. https://doi.org/10.18653/v1/P19-1050
Poria S, Majumder N, Hazarika D, Ghosal D, Bhardwaj R, Jian SYB, Hong P, Ghosh R, Roy A, Chhaya N, Gelbukh A, Mihalcea R (2021) Recognizing emotion cause in conversations. Cogn Comput 13:1317–1332. https://doi.org/10.1007/s12559-021-09925-7
Qu C, Che X, Yang Y, Zhang Z, Chang E, Zhang J, Zhu H, Yang L (2025) Enhancing emotion recognition in virtual reality: a multimodal dataset and a temporal emotion detector. Front Psychol. https://doi.org/10.3389/fpsyg.2025.1709943
Rahwan I, Cebrian M, Obradovich N, Bongard J, Bonnefon J-F, Breazeal C, Crandall JW, Christakis NA, Couzin ID, Jackson MO, Jennings NR, Kamar E, Kloumann IM, Larochelle H, Lazer D, McElreath R, Mislove A, Parkes DC, Pentland A, Roberts ME, Shariff A, Tenenbaum JB, Wellman M (2019) Machine behaviour. Nature 568:477–486. https://doi.org/10.1038/s41586-019-1138-y
Rashkin H, Smith EM, Li M, Boureau Y-L (2019) Towards empathetic open-domain conversation models: a new benchmark and dataset. In: Proceedings of the 57th annual meeting of the association for computational linguistics, pp 5370–5381. https://doi.org/10.18653/v1/P19-1534
Rhue L (2018) Racial influence on automated perceptions of emotions. SSRN working paper. https://doi.org/10.2139/ssrn.3281765
Riek LD (2012) Wizard of Oz studies in HRI: a systematic review and new reporting guidelines. J Hum Robot Interact 1(1):119–136
Ringeval F, Sonderegger A, Sauer J, Lalanne D (2013) Introducing the recola multimodal corpus of remote collaborative and affective interactions. In: Proceedings of the 10th IEEE international conference and workshops on automatic face and gesture recognition, pp 1–8. https://doi.org/10.1109/FG.2013.6553805
Roller S, Dinan E, Goyal N, Ju D, Williamson M, Liu Y, Xu J, Ott M, Smith EM, Boureau Y-L, Weston J (2021) Recipes for building an open-domain chatbot. In: Proceedings of the 16th conference of the European chapter of the association for computational linguistics, pp 300–325. https://doi.org/10.18653/v1/2021.eacl-main.24
Russell JA (1980) A circumplex model of affect. J Pers Soc Psychol 39(6):1161–1178. https://doi.org/10.1037/h0077714
Sabour S, Liu S, Zhang Z, Liu J, Zhou J, Sunaryo A, Lee T, Mihalcea R, Huang M (2024) EmoBench: evaluating the emotional intelligence of large language models. In: Proceedings of the 62nd annual meeting of the association for computational linguistics, pp 5986–6004. https://doi.org/10.18653/v1/2024.acl-long.326
Sambasivan N, Kapania S, Highfill H, Akrong D, Paritosh P, Aroyo LM (2021) “Everyone wants to do the model work, not the data work”: data cascades in high-stakes AI. In: Proceedings of the 2021 CHI conference on human factors in computing systems, pp 1–15. https://doi.org/10.1145/3411764.3445518
Scassellati B, Admoni H, Matarić M (2012) Robots for use in autism research. Annu Rev Biomed Eng 14:275–294. https://doi.org/10.1146/annurev-bioeng-071811-150036
Scherer KR (2001) Appraisal considered as a process of multilevel sequential checking. In: Appraisal processes in emotion: theory, methods, research. Oxford University Press, pp 92–120
Scherer KR (2005) What are emotions? And how can they be measured? Soc Sci Inf 44(4):695–729. https://doi.org/10.1177/053901840505821
Scherer KR, Moors A (2019) The emotion process: event appraisal and component differentiation. Annu Rev Psychol 70:719–745. https://doi.org/10.1146/annurev-psych-122216-011854
Schlicher M, Li Y, Murthy SMK, Sun Q, Schuller BW (2025) Emotionally adaptive support: a narrative review of affective computing for mental health. Front Dig Health. https://doi.org/10.3389/fdgth.2025.1657031
Schmidt P, Reiss A, Duerichen R, Marberger C, Van Laerhoven K (2018) Introducing WESAD, a multimodal dataset for wearable stress and affect detection. In: Proceedings of the 20th ACM international conference on multimodal interaction, pp 400–408. https://doi.org/10.1145/3242969.3242985
Searle JR (1980) Minds, brains, and programs. Behav Brain Sci 3(3):417–424. https://doi.org/10.1017/S0140525X00005756
Sharma A, Miner A, Atkins D, Althoff T (2020) A computational approach to understanding empathy expressed in text-based mental health support. In: Proceedings of the 2020 conference on empirical methods in natural language processing, pp 5263–5276. https://doi.org/10.18653/v1/2020.emnlp-main.425
Sharma A, Lin IW, Miner AS, Atkins DC, Althoff T (2021) Towards facilitating empathic conversations in online mental health support: a reinforcement learning approach. In: Proceedings of the web conference 2021. https://doi.org/10.1145/3442381.3450097
Shi Y, Yu K, Dong Y, Chen F (2026) Large language models in education: a systematic review of empirical applications, benefits, and challenges. Comput Educ Artif Intell 10:100529. https://doi.org/10.1016/j.caeai.2025.100529
Shingjergji K, Iren D, Urlings C, Klemke R (2026) Affective computing in online higher education: a systematic literature review. Comput Educ Artif Intell 10:100499. https://doi.org/10.1016/j.caeai.2025.100499
Shum H-Y, He X-D, Li D (2018) From Eliza to Xiaoice: challenges and opportunities with social chatbots. Front Inf Technol Electron Eng 19(1):10–26. https://doi.org/10.1631/FITEE.1700826
Slater M (2009) Place illusion and plausibility can lead to realistic behaviour in an immersive virtual environment. Philos Trans R Soc B 364(1535):3549–3557. https://doi.org/10.1098/rstb.2009.0138
Socher R, Perelygin A, Wu J, Chuang J, Manning CD, Ng AY, Potts C (2013) Recursive deep models for semantic compositionality over a sentiment treebank. In: Proceedings of the 2013 conference on empirical methods in natural language processing, pp 1631–1642
Sofroniew N, Kauvar I, Saunders W, Chen R, Henighan T, Hydrie S, Citro C, Pearce A, Tarng J, Gurnee W, Batson J, Zimmerman S, Rivoire K, Fish K, Olah C, Lindsey J (2026) Emotion concepts and their function in a large language model. Anthoropic
Somarathna R, Bednarz T, Mohammadi G (2023) Virtual reality for emotion elicitation—a review. IEEE Trans Affect Comput 14(4):2626–2645. https://doi.org/10.1109/TAFFC.2022.3181053
Spitale M, Axelsson M, Gunes H (2024) Appropriateness of LLM-equipped robotic well-being coach language in the workplace: a qualitative pilot study. https://doi.org/10.48550/arXiv.2401.14935 arXiv:2401.14935
Stark L (2018) Algorithmic psychometrics and the scalable subject. Soc Stud Sci. https://doi.org/10.1177/03063127187720
Stark L, Hoey J (2021) The ethics of emotion in artificial intelligence systems. In: Proceedings of the 2021 ACM conference on fairness, accountability, and transparency, pp 782–793. https://doi.org/10.1145/3442188.3445939
Stieglitz S, Mirbabaie M, Ross B, Neuberger C (2018) Social media analytics: challenges in topic discovery, data collection, and data preparation. Int J Inf Manag 39:156–168. https://doi.org/10.1016/j.ijinfomgt.2017.12.002
Tang C, Yu W, Sun G, Chen X, Tan T, Li W, Lu L, Ma Z, Zhang C (2024) SALMONN: towards generic hearing abilities for large language models. In: Proceedings of the international conference on learning representations. https://doi.org/10.48550/arXiv.2310.13289
Tian Y, Huang T, Liu M, Jiang D, Spangher A, Chen M, May J, Peng N (2024a) Are large language models capable of generating human-level narratives? In: Proceedings of the 2024 conference on empirical methods in natural language processing, pp 17659–17681. https://doi.org/10.18653/v1/2024.emnlp-main.978
Tian L, Wang Q, Zhang B, Bo L (2024b) EMO: emote portrait alive generating expressive portrait videos with Audio2Video diffusion model under weak conditions. In: Computer vision—ECCV 2024, pp 244–260. https://doi.org/10.1007/978-3-031-73010-8_15
Troiano E, Oberländer L, Klinger R (2023) Dimensional modeling of emotions in text with appraisal theories: corpus creation, annotation reliability, and prediction. Comput Linguist 49(1):1–72. https://doi.org/10.1162/coli_a_00461
Tsai Y-HH, Bai S, Liang PP, Kolter JZ, Morency L-P, Salakhutdinov R (2019) Multimodal transformer for unaligned multimodal language sequences. In: Proceedings of the 57th annual meeting of the association for computational linguistics, pp 6558–6569. https://doi.org/10.18653/v1/P19-1656
Turing AM (1950) Computing machinery and intelligence. Mind 49:433–460
Tzirakis P, Trigeorgis G, Nicolaou MA, Schuller BW, Zafeiriou S (2017) End-to-end multimodal emotion recognition using deep neural networks. IEEE J Sel Top Signal Proces 11(8):1301–1309. https://doi.org/10.1109/JSTSP.2017.2764438
Vaidyam AN, Wisniewski H, Halamka JD, Kashavan MS, Torous JB (2019) Chatbots and conversational agents in mental health: a review of the psychiatric landscape. Can J Psychiatry. https://doi.org/10.1177/0706743719828977
Veale M, Borgesius ZF (2021) Demystifying the draft EU Artificial Intelligence Act. Comput Law Rev Int 22(4):97–112. https://doi.org/10.9785/cri-2021-220402
Wang Y, Guo J, Bai J, Yu R, He T, Tan X, Sun X, Bian J (2025) InstructAvatar: text-guided emotion and motion control for avatar generation. In: Proceedings of the AAAI conference on artificial intelligence, vol 39, pp 8132–8140. https://doi.org/10.1609/aaai.v39i8.32877
Wankhade M, Kulkarni C, Rao ACS (2025) A survey on aspect base sentiment analysis methods and challenges. Appl Soft Comput 167(A):112249. https://doi.org/10.1016/j.asoc.2024.112249
Winner L (1980) Do artifacts have politics? Daedalus 109(1):121–136
Witte BD, Reynaert V, Kieken D, Jabbour J, Demarey C, Dumoulin A, Possik J (2026) Immersive virtual reality learning and cognitive load: a multiple-day field study. Comput Hum Behav 176:108853. https://doi.org/10.1016/j.chb.2025.108853
Xia R, Ding Z (2019) Emotion-cause pair extraction: a new task to emotion analysis in texts. In: Proceedings of the 57th annual meeting of the association for computational linguistics, pp 1003–1012. https://doi.org/10.18653/v1/P19-1096
Xu S, Chen G, Guo Y-X, Yang J, Li C, Zang Z, Zhang Y, Tong X, Guo B (2024) VASA-1: lifelike audio-driven talking faces generated in real time. In: Proceedings of the 38th international conference on neural information processing systems, pp 660–684. https://doi.org/10.5555/3737916.3737937
Yan W-J, Li X, Wang S-J, Zhao G, Liu Y-J, Chen Y-H, Fu X (2014) Casme II: an improved spontaneous micro-expression database and the baseline evaluation. PLoS ONE 9(1):86041. https://doi.org/10.1371/journal.pone.0086041
Yan HY, Morrow G, Yang K-C, Wihbey J (2025) The origin of public concerns over AI supercharging misinformation in the 2024 U.S. presidential election. In: Harvard Kennedy School misinformation review. https://doi.org/10.37016/mr-2020-171
Yannakakis GN, Spronck P, Loiacono D, André E (2013) Player modeling. In: Artificial and computational intelligence in games, vol 6, pp 45–59. https://doi.org/10.4230/DFU.Vol6.12191.45
Zadeh A, Zellers R, Pincus E, Morency L-P (2016) Multimodal sentiment intensity analysis in videos: facial gestures and verbal messages (CMU-MOSI). IEEE Intell Syst 31:82–88. https://doi.org/10.1109/MIS.2016.94
Zadeh A, Chen M, Poria S, Cambria E, Morency L-P (2017) Tensor fusion network for multimodal sentiment analysis. In: Proceedings of the 2017 conference on empirical methods in natural language processing, pp 1103–1114. https://doi.org/10.18653/v1/D17-1115
Zadeh AB, Liang PP, Poria S, Cambria E, Morency L-P (2018) Multimodal language analysis in the wild: CMU-MOSEI dataset and interpretable dynamic fusion graph. In: Proceedings of the 56th annual meeting of the association for computational linguistics. pp 2236–2246. https://doi.org/10.18653/v1/P18-1208
Zhang Y, Wang M, Wu Y, Tiwari P, Li Q, Wang B, Qin J (2024) DialogueLLM: context and emotion knowledge-tuned large language models for emotion recognition in conversations. Neural Netw 192:107901. https://doi.org/10.1016/j.neunet.2025.107901
Zhang W, Deng Y, Liu B, Pan S, Bing L (2024b) Sentiment analysis in the era of large language models: a reality check. In: Findings of the association for computational linguistics: NAACL 2024, pp 3881–3906. https://doi.org/10.18653/v1/2024.findings-naacl.246
Zhang Y, Zhao D, Hancock JT, Kraut R, Yang D (2025) The rise of AI companions: how human-chatbot relationships influence well-being. https://doi.org/10.48550/arXiv.2506.12605
Zhang Y, Yang X, Xu X, Gao Z, Huang Y, Mu S, Feng S, Wang D, Zhang Y, Song K, Yu G (2026) Affective computing in the era of large language models: a survey from the NLP perspective. Knowl-Based Syst 337:115411. https://doi.org/10.1016/j.knosys.2026.115411
Zhao S, Wang S, Soleymani M, Joshi D, Ji Q (2019) Affective computing for large-scale heterogeneous multimedia data: a survey. ACM Trans Multimed Comput Commun Appl 15(93):1–32. https://doi.org/10.1145/3363560
Zhao J, Zhang T, Hu J, Liu Y, Jin Q, Wang X, Li H (2022) M3ED: multi-modal multi-scene multi-label emotional dialogue database. In: Proceedings of the 60th annual meeting of the association for computational linguistics. Association for Computational Linguistics, pp 5699–5710. https://doi.org/10.18653/v1/2022.acl-long.391
Zheng W-L, Lu B-L (2015) Investigating critical frequency bands and channels for EEG-based emotion recognition with deep neural networks. IEEE Trans Auton Ment Dev 7(3):162–175
Zheng C, Sabour S, Wen J, Zhang Z, Huang M (2023) AugESC: dialogue augmentation with large language models for emotional support conversation. In: Findings of the association for computational linguistics: ACL 2023. Association for Computational Linguistics, pp 1552–1568. https://doi.org/10.18653/v1/2023.findings-acl.99
Zhou L, Gao J, Li D, Shum H-Y (2018) The design and implementation of XiaoIce, an empathetic social chatbot. Comput Linguist 46(1):53–93. https://doi.org/10.1162/coli_a_00368
Ziems C, Held W, Shaikh O, Chen J, Zhang Z, Yang D (2024) Can large language models transform computational social science? Comput Linguist 50:237–291. https://doi.org/10.1162/coli_a_00502
Zuboff S (2019) The age of surveillance capitalism: the fight for a human future at the new frontier of power. PublicAffairs
Acknowledgements
The author gratefully acknowledges the support of JST SPRING, the Nippon Foundation HUMAI Program, the GMO Internet Foundation, and the Telecommunications Advancement Foundation, which provided the funding and research environment that enabled this work.
Funding
This work was supported in part by JST SPRING under Grant Number JPMJSP2150, the Nippon Foundation HUMAI Program, the GMO Internet Foundation, and the Telecommunications Advancement Foundation.
Author information
Authors and Affiliations
Contributions
The author confirms sole responsibility for the following: study conception and design, methodology, analysis and interpretation of the results, and manuscript preparation.
Corresponding author
Ethics declarations
Conflict of interest
The authors declare no competing interests.
Ethics approval and consent to participate
Not applicable.
Consent for publication
Not applicable.
Use of AI tools
The author used AI-based language tools to support translation and language editing during manuscript preparation. All AI-assisted outputs were reviewed, revised, and verified by the author, who takes full responsibility for the final content of the manuscript.
Additional information
Publisher's Note
Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
Appendix A: Coding rubric for the prior-surveys comparison table
Appendix A: Coding rubric for the prior-surveys comparison table
This appendix documents the rubric used to score each prior survey listed in Table 1 along the two diagnostic axes used in this paper: cross-layer dependence treated as central and prescriptive stance. The purpose of the rubric is not to rank prior surveys but to make the differentiation claim auditable: a future reader who disagrees with the placement of any row should be able to identify which scoring anchor they would invoke instead and why.
1.1 A.1 Scoring axes
Axis 1: Cross-layer dependence treated as central. This axis asks whether the survey treats inter-layer dependence—for instance theory–data, modeling–interaction, technology–ethics, or data–application—as a central analytic object rather than as a residual concern. The three-level scale is:
-
Yes: the survey thematically (at the section or chapter level) frames inter-layer disconnections as the object of analysis, and the structure of the survey is organized around such disconnections rather than around a single layer or a single modality.
-
Partial: the survey discusses two or more layers in interaction (e.g., couples ethics to method, or data to deployment), but the inter-layer disconnection itself is not the principal organizing object; cross-layer concern surfaces locally rather than systemically.
-
No: the survey is organized within a single layer (e.g., data, modeling, modality), reviews multiple layers in parallel without analyzing their dependence, or adopts a bibliometric stance in which layer-level analytic structure is not used.
Axis 2: Prescriptive stance. This axis asks whether the survey advances a prescriptive program that goes beyond taxonomy or empirical taking-stock and is intended to constrain subsequent research practice. The two-level scale used in Table 1 is:
-
Yes: the survey advances a normative program, such as a foundational vision for the field, a prescriptive ethics agenda, or a set of design recommendations intended to constrain how subsequent work should proceed. The recommendation may be confined to one layer (e.g., ethics) or may be field-wide; what matters is that the survey takes a position rather than merely describing the literature.
-
No: the survey is descriptive, taxonomic, or bibliometric in stance. Local recommendations may appear (most surveys close with “future directions”), but the survey does not commit to a programmatic position whose acceptance would reshape research practice.
A finer-grained four-level scale (Strong/Moderate/Weak/None) was considered during rubric development; it was collapsed to the binary Yes/No scale used in Table 1 because the disambiguation required to separate Strong from Moderate could not be applied uniformly across the eleven surveys without introducing reviewer-dependent variance. The collapse is conservative: surveys scored Yes satisfy at least the Moderate threshold, and surveys scored No do not reach the Moderate threshold. The richer scale is documented here for transparency but is not used in the table itself.
1.2 A.2 Inclusion criteria for the comparison set
The comparison set in Table 1 consists of eleven prior surveys selected under the following inclusion rules.
-
1.
Surveys, not primary papers. Only works whose self-stated genre is review, survey, manifesto, or ethics sheet are included. Influential primary papers cited elsewhere in this paper (e.g., Ekman 1992, Russell 1980, Barrett 2017) are not eligible for the differentiation table because their genre is theoretical or empirical contribution rather than survey.
-
2.
Direct relevance to affective computing as a field. Surveys whose object is affective computing, emotion recognition, affect detection, emotion in dialog, or the ethics of emotion AI are included. Surveys of adjacent fields (general HCI, general dialog systems, general AI ethics) are excluded unless their scope explicitly subsumes affective computing.
-
3.
Temporal coverage. The set spans 1997–2025 and is intentionally chosen to include the foundational manifesto (Picard 1997), mature interdisciplinary reviews from the 2010s (Calvo and Mello 2010; Mccoll et al. 2016; Poria et al. 2017, 2019a; Zhao et al. 2019), the inflection point of ethics-explicit work (Mohammad 2022), and the 2024–2025 LLM- and foundation-model-era surveys (Pei et al. 2024; Kumar et al. 2025; Moon et al. 2026; Zhang et al. 2026). The intent is to cover the inflection points of the field rather than to be exhaustive.
-
4.
Representative scope. Where multiple surveys cover essentially the same scope in the same year, the comparison set retains the one most cited at the time of writing or, when citation counts are similar, the one whose framing is closest to the cross-layer concern of this paper. The two Dey et al. 2025 surveys are both retained because they target distinct slices (general multimodal MER vs. LLM-/foundation-model-based MER 2021–2025) and the differentiation between them is itself diagnostic of the field’s current trajectory.
-
5.
Exclusions. Workshop-only proceedings, vendor white papers, and surveys whose published version was not available at the time of writing are excluded. This rubric does not claim that no other survey could have been included; it claims that the eleven included surveys are representative of the comparison categories the differentiation table is designed to span.
1.3 A.3 Disambiguation procedures for borderline cases
Three disambiguation procedures were applied during scoring; they are documented here so that the borderline judgments are auditable.
D1: Foundational vision vs. cross-layer diagnosis. A survey whose mission is to define a new field (e.g., Picard 1997) is necessarily multi-topic: data, models, applications, and ethics are all touched. The rubric does not score such breadth as Yes on Axis 1. Yes on Axis 1 requires that the inter-layer dependence (one layer’s choice constraining another layer’s evaluation) be the analytic object. Foundational manifestos that span layers but do not analyze their dependence as such are scored No on Axis 1 even when they take a strong programmatic position on Axis 2.
D2: Ethics-and-method coupling. A survey whose argument explicitly couples ethical evaluation to methodological design (e.g., Mohammad 2022, whose 50 considerations span task design, data, method, evaluation, and implications) is scored Partial on Axis 1. The coupling is real and intentional, but it does not reach the full cross-layer-as-central threshold: theory–data and modeling–interaction couplings are not the principal analytic objects. We note that the method layer (L3) and the ethics layer (L6) are not adjacent strata in the six-layer pipeline of Sect. 1; an ethics-method coupling is therefore substantively an L3–L6 inter-layer coupling, which is precisely why it justifies a Partial score rather than No, even though it falls short of the systemic cross-layer-as-central threshold required for Yes. Partial is also used when a survey integrates two adjacent layers without taking the overall layered structure as its frame.
D3: Bibliometric and taxonomic surveys. A survey whose primary method is bibliometric (citation analysis, co-occurrence mapping; e.g., Pei et al. (2024)) is scored No on Axis 1 even when its descriptive coverage formally spans every layer. The rationale is that bibliometric methods describe the distribution of attention across topics rather than analyze inter-layer dependence; without an analytic claim about dependence, parallel coverage is not cross-layer dependence in the sense this paper requires.
1.4 A.4 Per-survey scoring rationales
The following entries provide, for each row of Table 1, the scoring rationale on both axes. The scores reproduced below match Table 1 cell-for-cell; any divergence should be treated as an error. Where the rationale relies on disambiguation procedures D1–D3, the procedure is named.
1. Picard (1997). Axis 1: No. Foundational manifesto for the field; layers (data, models, applications, ethics) are surveyed but inter-layer dependence is not the analytic object (D1). Axis 2: Yes (foundational vision, not cross-layer diagnosis). The work issues a programmatic position by defining what affective computing should be. It is scored Yes on Axis 2 for that reason. The qualifier in Table 1 explicitly distinguishes this stance from cross-layer diagnosis.
2. Calvo and Mello (2010). Axis 1: No (layers reviewed in parallel). The survey covers face, voice, body, physiology, EEG, text, and multimodal channels in parallel, alongside emotion theory; coverage is broad, but the dependence among these strata is not made into the analytic object. Axis 2: No. The work is an interdisciplinary review with descriptive intent; it does not advance a prescriptive program in the sense Axis 2 requires.
3. Mccoll et al. (2016). Axis 1: No. The scope is autonomous human affect detection methods within HRI scenarios; the focus is methodological taxonomy within a single application context rather than inter-layer dependence. Axis 2: No. The survey is descriptive, oriented toward HRI practitioners selecting detection methods; no field-wide prescriptive position is advanced.
4. Poria et al. (2017). Axis 1: No (modality-centric). The survey is organized around unimodal (visual, audio, text) analysis and multimodal fusion; modality is the organizing principle, and inter-layer dependence is not foregrounded. Axis 2: No. Comprehensive review aimed at modeling-layer practitioners; the closing remarks discuss future directions but do not amount to a prescriptive program over the field.
5. Poria et al. (2019a). Axis 1: No (task-internal). The scope is the ERC task: research challenges, datasets, and recent neural approaches, such as CMN, ICON, and DialogueRNN. The analysis stays within one task family rather than treating cross-layer dependence as its central analytic object. Axis 2: No. Systematic review of one task; descriptive in stance.
6. Zhao et al. (2019). Axis 1: No. Focus on heterogeneous multimedia (images, music, videos, multimodal) at the modeling and data layers; cross-layer dependence is not the analytic object. Axis 2: No. Modality-and-content-centered survey without a prescriptive program.
7. Mohammad (2022). Axis 1: Partial (ethics–method coupling). The 50 considerations span task design, data, method, evaluation, and implications; this is a deliberate coupling of ethical reasoning to methodological choice (D2). The coupling does not extend to the full theory–data and modeling–interaction couplings that would warrant Yes, but it reaches further than No. Axis 2: Yes (prescriptive ethics agenda). The ethics sheet format itself is a programmatic intervention: it tells the field how ethical reasoning should be incorporated into emotion-recognition and sentiment-analysis research. This is a clear prescriptive program, even though its scope is bounded by the ethics layer.
8. Pei et al. (2024). Axis 1: No (bibliometric, not analytic). A bibliometric review of 33,448 papers from 1997–2023; the method is descriptive distribution analysis (D3). Even where every layer appears in the dataset, parallel coverage is not cross-layer dependence in the sense Axis 1 requires. Axis 2: No. The method does not yield a programmatic position; the contribution is a map of the literature, not a stance on it.
9. Kumar et al. (2025). Axis 1: No. Comprehensive survey of multimodal emotion recognition; modality and architectural lineage are the organizing axes. Axis 2: No. The survey is descriptive of recent technical advances; future-direction remarks do not constitute a prescriptive program.
10. Moon et al. (2026). Axis 1: No. Survey of multimodal emotion recognition 2021–2025 with focus on LLM- and foundation-model-based architectures; concentrates on the modeling layer with attention to architectural taxonomy of LLM-based MER. Axis 2: No. Technical review; descriptive rather than prescriptive in stance.
11. Zhang et al. (2026). Axis 1: No. LLM-era affective computing survey from an NLP perspective covering affect understanding and affect generation tasks, instruction tuning, prompt engineering, and benchmarks; the analysis stays at the modeling and benchmark layers. Axis 2: No. Technical survey oriented toward NLP practitioners; future-direction remarks are local rather than programmatic over the field.
1.5 A.5 Limitations of the rubric
Three limitations of the rubric are recorded here in the interest of fair self-assessment.
First, the scoring was performed by the authors of this paper without an independent second coder; inter-rater reliability (IRR) is therefore not measured. This is a real limitation, and the position-paper genre is not, in itself, an excuse for it. The decision to defer second-coder scoring rests on two grounds: (i) the binary collapse described in Sect. A.1 is intended precisely to reduce the surface on which IRR variance acts, and (ii) the per-survey rationales above expose every disambiguation step so that an independent reader can reproduce the scoring with public information. A future revision of this paper, or a follow-up, could readily fold in second-coder scoring; the rubric is designed to support that.
Second, the inclusion criteria of Sect. A.2 are themselves selection choices, and any selection choice is in principle contestable. The rubric’s defense is not that the eleven surveys are uniquely correct but that the comparison categories they span (foundational manifesto, interdisciplinary review, modality-centric review, task-centric review, ethics sheet, bibliometric review, LLM/foundation-model survey) are the categories against which a position-paper claim of novelty must be checked.
Third, the rubric does not score any survey on the further axis of cascade depth, that is, whether the survey, when it does treat inter-layer dependence, treats the dependence as a single coupling or as a chained cascade of couplings. Adding such an axis would be informative, but it would also collapse onto the position of this paper itself (since this paper is the explicit cascade claim), and the rubric’s purpose here is comparison, not self-promotion. The absence of this axis is therefore a deliberate scope decision rather than an oversight.
The rubric, the per-survey rationales, and the table together support the position-paper claim of Sect. 8 and the prescriptive program of Sect. 9 at the level of evidence: any reader who disagrees with the placement of a row in Table 1 is invited to re-score that row using this rubric and to re-examine, against the new score, whether the cascade diagnosis and the DC1–DC5 program continue to follow.
Rights and permissions
Springer Nature or its licensor (e.g. a society or other partner) holds exclusive rights to this article under a publishing agreement with the author(s) or other rightsholder(s); author self-archiving of the accepted manuscript version of this article is solely governed by the terms of such publishing agreement and applicable law.
About this article
Cite this article
Inoshita, K. Bridging the silos in affective AI: a critical perspective from data to society. AI & Soc (2026). https://doi.org/10.1007/s00146-026-03324-y
Received:
Accepted:
Published:
Version of record:
DOI: https://doi.org/10.1007/s00146-026-03324-y
