| dc.description.abstract |
Sign Language-based communication is a major source of communication between
people with hearing impairments and the general public. Hand, head, body, and gesture
developments are immediate aspects of communicating specific feelings or techniques of
correspondence in any language, whether marked or spoken. Hand sign recognition is
known as manual sign language recognition in sign language, whereas facial expressions
and body gesture understanding are known as non-manual sign language recognition.
Several researchers are focusing either on manual gestures or non-manual gestures
separately; a rare focus is on manual and non-manual gestures concurrently, making the
loss of content or complete meaning of the sentence. Nonetheless, one of the great
challenges of Sign Language Recognition is dealing with formal and non-formal
parameters simultaneously within a single framework or methodology whereas using
dynamic sign language recognition. However, dynamic sign language recognition has
several difficulties recognizing complex features, accurate classifications, and extensive
video sequence data training and testing in the Spatio-temporal domain. To issues, the
current research study presents a Multimodal Dynamic Sign Language Recognition in the
Spatio-temporal domain based on vision-based deep learning. This research consists of
two-fold contributions. First, for the training and testing, there is a lack of such a dataset,
where manual and non-manual modalities combine with affective facial expressions. So
first contribution is compiling a Pakistan Sign Language (PSL) dataset with Manual
and Non-Manual modalities named PkSLMNM. Further, we proposed a system called Sign
Language Action Transformer Network (SLATN) is restricts hand, body, and facial signals
in video arrangements. Here we are using a Transformer-style structural design as a "base
network" to remove highlights from a spatiotemporal space. The model hastily figures out
how to follow individual people and their setting of activity in different edges. Further, a
"head network" at the same time classify and region out hand movement and facial
expression, which is frequently critical to figuring out communication through signing. It
uses its attention mechanism for creating tight bounding boxes around classified gestures.
Later, the model's performance is evaluated against state-of-the-art datasets and
conventional identification techniques. It not only completes tasks more efficiently but also
with good accuracy. Our suggested network achieves 82.66% testing accuracy and 94.13
Giga FLOPs of a notable processing performance |
en_US |