Facial Expressions Recognition in Sign Language Based on a Two-Stream Swin Transformer Model Integrating RGB and Texture Map Images

Descripción del Articulo

The study of facial expressions in sign language has become a significant research area, as these expressions not only convey personal states, but also enhance the meaning of signs within specific contexts. The absence of facial expressions during communication can lead to misinterpretations, unders...

Descripción completa

Detalles Bibliográficos
Autores: Ramirez Cerna, Lourdes, Rodriguez Melquiades, Jose, Escobedo Cárdenas, Edwin Jhonatan, Cámara Chávez, Guillermo, Garcia Miranda, Dayse
Formato: artículo
Fecha de Publicación:2025
Institución:Universidad de Lima
Repositorio:ULIMA-Institucional
Lenguaje:inglés
OAI Identifier:oai:repositorio.ulima.edu.pe:20.500.12724/24467
Enlace del recurso:https://hdl.handle.net/20.500.12724/24467
https://doi.org/10.13053/CyS-29-2-5119
Nivel de acceso:acceso abierto
Materia:Pendiente
https://purl.org/pe-repo/ocde/ford#2.02.04
Descripción
Sumario:The study of facial expressions in sign language has become a significant research area, as these expressions not only convey personal states, but also enhance the meaning of signs within specific contexts. The absence of facial expressions during communication can lead to misinterpretations, underscoring the need for datasets that include facial expressions in sign language. To address this, we present the Facial-BSL dataset, which consists of videos capturing eight distinct facial expressions used in Brazilian Sign Language. Additionally, we propose a two-stream model designed to classify facial expressions in a sign language context. This model utilizes RGB images to capture local facial information and texture map images to record facial movements. We assessed the performance of several deep learning architectures within this two-stream framework, including Convolutional Neural Networks (CNNs) and Vision Transformers. In addition, experiments were conducted using public datasets such as CK+, KDEF-dyn, and LIBRAS. The two-stream architecture based on the Swin Transformer model demonstrated superior performance on the KDEF-dyn and LIBRAS datasets and achieved a second-place ranking on the CK+ dataset, with an accuracy of 97% and an F1-score of 95%.
Nota importante:
La información contenida en este registro es de entera responsabilidad de la institución que gestiona el repositorio institucional donde esta contenido este documento o set de datos. El CONCYTEC no se hace responsable por los contenidos (publicaciones y/o datos) accesibles a través del Repositorio Nacional Digital de Ciencia, Tecnología e Innovación de Acceso Abierto (ALICIA).