Sentiment Analysis of German Sign Language Fairy Tales
arXiv:2604.16138v1 [cs.CL] 17 Apr 2026
Fabrizio Nunnari, Siddhant Jain, Patrick Gebhard {fabrizio.nunnari, siddhant.jain, patrick.gebhard}@dfki.de German Research Center for Artificial Intelligence (DFKI) Saarland Informatics Campus D3 2, 66123 Saarbrücken, Germany April 20, 2026
Abstract
Although performance of SL translation pipelines is still rudimentary, some companies start offering online automated translation tools (e.g.: Signapse - https://www.signapse.ai, Migam - https://migam.ai, and Nagish https://sign.mt). The output of SL synthesis systems is however criticized as being too robotic and unexpressive, to the extent that the semantics of the message is not correctly communicated. One reason for that, is the lack of explicit support for the communication of “sentiments”, and “emotions” in SL translation systems. Expressing emotions in spoken communication (e.g., through facial expressions, hand gestures, and body postures) can be seen as an addition to the basic message conveyed via words. In contrast, in SL the whole body and face are already involved in the communication of the semantics of the message, and such a dichotomy between grammatical communication and emotion expression is not possible. Research on emotion expression in SL is still emerging (e.g., see the ExEmSiLa workshop [8]), but it is carried on mainly by SL linguists through time-consuming manual annotation and human observations of video material, without the support of modern automated machine learning tools. In contrast, for written languages, sentiment analysis is a well established branch of natural language processing (NLP) that automatizes the inference of an emotion (e.g., happiness, sadness) or the valence (positive, neutral, negative). This paper investigates the automated sentiment analysis of SL. The goal being to deploy a computational model that can correctly infer valence
We present a dataset and a model for sentiment analysis of German sign language (DGS) fairy tales. First, we perform sentiment analysis for three levels of valence (negative, neutral, positive) on German fairy tales text segments using four large language models (LLMs) and majority voting, reaching an inter-annotator agreement of 0.781 Krippendorff’s alpha. Second, we extract face and body motion features from each corresponding DGS video segment using MediaPipe. Finally, we train an explainable model (based on XGBoost) to predict negative, neutral or positive sentiment from video features. Results show an average balanced accuracy of 0.631. A thorough analysis of the most important features reveal that, in addition to eyebrows and mouth motion on the face, also the motion of hips, elbows, and shoulders considerably contribute in the discrimination of the conveyed sentiment, indicating an equal importance of face and body for sentiment communication in sign language. Keywords: German sign language, DGS, sentiment analysis, fairy tales, explainable AI, machine learning.
1
Introduction
Sign languages (SLs) are the native language of 70 million people in the world [6] and they are the main communication languages among Deaf communities. SLs rely on hands, fingers, torso, shoulders, gaze, head motion and facial expressions to convey meaning. 1
timent classification [10]. Four publicly available large language models (LLMs) are used to extract sentiment labels from the text segments, and a final label is assigned through majority voting. Second, the MediaPipe library [12] is used to analyse SL videos and extract blend-shapes and landmarks, which are further elaborated to extract information about velocities, accelerations, mutual distances, accumulated distances, bone rotations, and peak frequency. Finally, XGBoost is used to build a model predicting sentiment from video features. The details are reported in section 3. This paper brings the following three contributions. First, we release DGS-Fabeln-1-SE (Sentiment Estimation), which extends DGS-Fabeln-1 with: i) the sentimental labels predicted by the 4 LLMs and their majority voting, and ii) the MediaPipe motion data of all segments, including an extra set of features precomputed for the classification task. Second, we deploy an explainable prediction model, based on XGBoost, capable of predicting sentiment in DGS-Fabeln-1 video segments. Third, we report and comment on the list of the most important features for the classification of valence from SL videos. The DGS-Fabeln-1-SE dataset is publicly available at https://doi.org/ 10.5281/zenodo.18879038. The code to process the dataset is available at https://github.com/DFKI-SignLanguage/ LREC2026-DGS-Fabeln-1-SE
Figure 1: A frame from the first sentence of the tale The Hare and the Hedgehog from the DGS-Fabeln1 corpus. The corresponding German sentence is “Der Hase und der Igel. Es war einmal: So fangen Märchen an. Ein Märchen ist eine sehr alte Geschichte. Dieses Märchen heißt: Der Hase und der Igel. Das Märchen geht so:” [ENG: “The Hare and the Hedgehog. Once upon a time: that’s how fairy tales begin. A fairy tale is a very old story. This fairy tale is called: The Hare and the Hedgehog. The fairy tale goes like this:”]. The corresponding video lasts about 15 seconds. from the analysis of SL video segments1 . Valence (also known as Pleasure) is one of the three dimensions used to quantify emotions in the PAD space (Pleasure/Valence, Arousal, Dominance) [13]. In this first investigation, we target only the valence dimension. For this work, we rely on the DGS-Fabeln-1 dataset [16, 15], a parallel corpus of 7 German sign language fairy tales, consisting of 574 German text segments each associated with a video clip. Figure 1 shows an example. We focus on fairy tales because they naturally cover the whole range of positive to negative valence, thus displaying a high degree of expressivity in sign language performance. This seems to be a better candidate with respect to the most popular dataset for DGS used in natural language processing, i.e., the RWTH-PHOENIX-2014-T corpus [2], which is limited to the domain of weather forecast. To perform our tests, we first take advantage of the state-of-the-art performance of LLMs in sen-
2
Related work
2.1
Sentiment LLMs
analysis
through
Manual sentiment annotation is a time- and resource-consuming task. Given the increased reliability of automated sentiment analysis, we opted for a procedural annotation of the DGS-Fabeln-1 dataset. The system introduced by [7] is considered among the best for sentiment analysis on the German language. However, an early test on fairy tales showed a peculiar behavior, with more than 90% of the text chunks classified as “neutral”. A manual 1 We refer to a SL video clip conveying a semantic information as a segment, instead of sentence, because there is check on a few samples led us to consider such lino exact correspondence between the two. brary inappropriate for this task. 2
Indeed, Klähn et al. [10] report that while finetuned German BERT models achieve 94-96% F1 scores on contemporary German text, these models experience catastrophic performance degradation in other contexts, like historical German. In contrast, LLMs can generalize much better in unseen domains (zero-shot learning). For example, also Suter and Mecker [20] report that GPT 4 demonstrated a “substantial to near perfect agreement” with human annotators on German text classification. In our work, we relied on a combination of the latest available versions of four different LLMs (GPT5, Sonic, Mistral, GPT OSS 20B) and ran a majority voting to compensate for disagreements.
2.2
50 sentences in Chinese sign language from 18 participants under 5 emotions. The emotions were derived from a clusterization of 5 areas of the Ushape typically formed by correlating Valence and Arousal classification on sentiments. They claim 88% accuracy in the classification task. Their system, based on wearable armband with builtin surface electromyograph (sEMG) and inertial measurement unit (IMU) sensors is unpractical for every-day use, but their results confirm that a significant part of the emotions of a message is carried by body movement. [5] introduced recently a dataset of 200 sentences for sentiment analysis on American sign language (ASL). Each video was annotated by three SL experts for sentiment valence, 10 emotions, and free comments. The creation process represents the state-of-the-art in the field, but the resulting annotations present an averge inter-annotator agreement on only 0.53 in terms of Krippendorff’s Alpha, with a couple of labels scoring less than 0.2. This highlights the difficulty and subjectivity of human annotation for SL videos. Differently, in our work we use automated annotation only on the textual side of the corpus: a task for which LLMs show a satisfying accuracy and higher agreement. [5] also presented an attempt to automatize the prediction of emotions from videos. They based their work on multimodal LLMs with video analysis capabilities, trying to predict emotion labels from an encoding of the video downsampled at 10Hz rate. The prediction of 3-levels of sentiment reaches 56.18% accuracy using AffectGPT, but with the need of high computation power and without an explanation of the result. In contrast, in this work we employ a light-weight preprocessing phase for body/face motion features extraction, followed by a very light-weight XGBoost predictor (possibly enabling real-time inference) which gives by-product a list of the most relevant features.
Sentiment Analysis on SLs
[17] performed sentiment analysis on Indian sign language. They trained a mixed model using both video and text to predict labels in 4193 frames, singularly annotated for negative, neutral, and positive valence. They claim that, using VGG16, the video-only modality reaches >99% accuracy. In contrast, we performed the analysis on full sentences, rather than single frames, and used an explainable model listing the human-understandable features that contribute to the classification. [21] investigated a 3-level sentiment analysis for single signs of the Turkish sign language (TİD). They used a late fusion architecture merging the classification based on face and hands separately, reporting that the hands-only modality reaches the worst accuracy (42%) with respect to using faceonly (46%) or their combination (65%). The main difference from our work is that they perform the classification on short single-sign videos, while we classify full sentences. [23] implemented a system that converts sign language into emotionally driven speech. To achieve it, they first analyse facial expressions from SL videos and use the derived pleasure-arousal-dominance (PAD) values to drive the speech synthesis. Results show a good match between the emotions recognized in the facial expressions w.r.t. speech, but this doesn’t assess the validity of using PAD recognition as a good indicator of emotion in SL, as the face is simultaneously involved in performing grammar roles. [24] built a neural model tested on a dataset of
2.3
Sentiment Tales
Analysis
on
Fairy
There are several reasons on why to prefer fairy tales (over other domains such as weather forecast) to perform sentiment analysis. First, fairy tales display much larger standard deviations across all emotion categories, indicating 3
3.2
greater variability in emotional content within individual tales [14].
Method overview
Figure 2 shows the steps for the preparation of the training/testing data. As depicted on the left side, the text segments of each fairy tale were given as input to four LLMs, encapsulated in a prompt asking to perform sentiment analysis separately on each segment, and also indicating whether more sentiments are applicable. One single label per segment is then chosen via majority voting. Text segments presenting mixed emotions or no agreement were discarded. On the right side, Figure 2 depicts the video processing steps. Each video segment is processed with the MediaPipe library [12]. The raw 3D landmark locations and facial blendshapes are used to compute an extra set of manually engineered features, and then aggregated at segment level via means and standard deviations. Predicted sentiments and video aggregated features are then joined into the DGS-Fabeln-1-SE dataset. Because of the absence of some videos in the original DGS-Fabeln-1 corpus, only 517 segments are finally available.
[22] report that fairy tales present distinctive challenges and opportunities for sentiment analysis. The genre’s formulaic nature, with clear archetypal characters and predictable narrative structures, provides well-defined emotional contexts that facilitate automatic sentiment detection. In addition, [22] have identified specific emotional trajectory patterns in fairy tales, with polynomial fitting revealing wave-shaped emotional flows common across different tales. Coherently, [14] reports that the emotional dynamics follow predictable patterns from initial states through conflict development to resolution, often culminating in positive outcomes that justify the “happily ever after” convention. We will use those observations to verify the plausibility of LLM-base automated sentiment prediction on our fairy tales text. Finally, it seems that the high emotional variability within individual tales suggests complex emotional journeys that engage readers through dynamic sentiment shifts rather than static positive messaging [14, 9]. This is in accordance to our first observations that valence may change rapidly even within a single video segment. For this reason, as described later in the method section, we implemented a procedure to detect and exclude segments with mixed emotions.
3.3
Feature extraction from videos
Each frontal video segment from the DGS-Fabeln-1 dataset was processed using a multimodal feature extraction pipeline based on MediaPipe Holistic and Face Landmarker v2 models [12]. The aim was to obtain consistent three-dimensional trajectories of upper body, hands, head, and facial blendshapes for every frame. The following landmark sets were extracted: Pose landmarks for wrists, torso, shoul3 Method ders, elbows, hips, and nose; Facial blendshapes and transformation matrix for 52 blendshape activations and a 4×4 affine transform containing head 3.1 Data: DGS-Fabeln-1 rotation and translation. Because landmark detection occasionally fails For our data preparation, we use the DGS-Fabeln(e.g., during occlusion or motion blur), missing val1 [16, 15], a dataset of seven German fairy tales ues were linearly interpolated over time to ensure interpreted in German sign language (DGS) from frame continuity. simplified German text. Each fairy tale is divided From the interpolated time series, additional feainto text segments (between 37 and 150) for a totures, commonly used in gesture recognition, were tal of 574 segments. Each segment, roughly corcomputed to better describe the motion dynamics. responding to 2.5 sentences, is associated with a video clip. Video clips show the DGS interpreta• Velocity and acceleration for wrists, shoultion of the corresponding text, each lasting 9.6 secders, hips, and nose. onds on average. The dataset totals 92 minutes of video recording. Figure 1 shows an example. • Euclidean distances between: i) the two 4
Classify Sentiment 574 text segments
GPT5
(via Perplexity)
Filter out mixedsentiment samples
Sonic
(via Perplexity)
Join Majority Voting
Segment aggregation (means, SDs)
517 segments
Gemini Flash 2.5 Sentiment Extraction Prompt
distance, velocity, acceleration
520 segments
Train/Test Dataset GPTOSS 20B
Extra features computation
Media Pipe processing
landmarks and blendshapes
tale, segment, text, sentiment, body/face motion features (x 288) …
Figure 2: Overview of our dataset preparation pipeline. wrists, ii) the two elbows iii) a wrist and its GPT5, Sonic, Gemini Flash 2.5, and GPT OSS corresponding shoulder, iv) a wrist and the op- 20B. For GPT5 and Sonic, the analysis was perposite shoulder, v) wrists and the nose. formed through the Perplexity Pro app. Gemini was used through its web interface. Finally, GPT • Head pitch, yaw, and roll angles computed OSS 20B was run on a local machine using the olwith a decomposition of the 3×3 rotation sub- lama3 service and its Python API. matrix of the head transform. All sentiment voting was performed using the prompt reported in Listing 1, followed by the tale • Left and right arm elevation angles comsentences, each surrounded by double quotes and puted from the shoulder–elbow vectors relative separated by an extra space. to the torso axis. The goal of the prompt is to perform the senFor each video, all of those frame-level features timent judgment and to spot if mixed sentiments were aggregated into the following statistics: appear in the same text chunk. Our goal was to skip such segments in order to avoid the treatment • Mean. of a multi-label problem. The resulting votes were then aggregated into a • Standard deviation. single column through majority voting. The major• Accumulated movement distance for ity agreement on a segment is reached if a label has wrists, shoulders, and nose. more counts than any other. However, a segment was marked as “no-agreement” if either: • Peaks per second calculated on each raw data: useful to catch signal sparks that might • Two (or more) labels have the same count, or be hidden by average aggregations over long sequences. Peaks were computed via the • the majority label is itself a mixed emotion. find peak function of the SpiPy framework2 and then normalized by video duration. Figure 3 shows the final distribution of the labels after aggregation. In total, 520 segments passed the This yielded to a feature vector of 396 items per majority voting filter. video segment. As a measure of inter-annotator agreement, we computed the Krippendorff’s alpha coefficient [11], 3.4 Sentiment extraction from text using the Castro’s implementation for Python [3]. Table 1 shows the alpha values of the collected votes The sentiment analysis was performed by aggrebefore and after applying the majority voting filter. gating the votes of four large language models: As expected, the alpha coefficient increases (from 2 https://docs.scipy.org/doc/scipy/reference/ generated/scipy.signal.find_peaks.html
3 https://ollama.com
5
Listing 1: Sentiment analysis prompt Consider the task of Sentiment analysis ( also called opinion mining ) in the field of Natural Language Processing ( NLP ). I will give you a list of German text chunks . For each chunk , evaluate its " sentiment " on a 3 levels scale : negative , neutral , positive . Also , for each chunk , mark if there is more than one sentiment that can be applied ( for example , if a long text chunk changes sentiment ) , and report the sentiments separated by a dash ( -) Format the output as a three - column comma - separated values ( CSV ) , using the comma as separator . The first column will contain the original text chunk . The second column will contain the ( list of ) sentiment ( s ) that you evaluated . The third column will contain " yes " if multiple sentiments were found , " no " otherwise . Also , output a header with column names : " Text " , " Sentiments " , " Multi ". As a consequence , the number of output rows must be the same as the number of input text chunks . In the output , omit any kind of preamble and explanation . Output only the CSV data . Do not use double quotes in the output CSV . Here is the list of chunks :
Finally, for a comparison with other sentiment 0.715 to 0.786) after filtering out the items with no classification tasks, we computed a Pearson ρ corremajority applicable. lation after mapping negative, neutral, and positive Data Samples Krippendorff’s alpha labels to −1, 0, and +1, respectively. Raw votes 574 0.715 All metrics were calculated using the scikit-learn After majority voting 520 0.786 library [18]. The macro-averaged F1-score was used as the primary criterion for model selection during Table 1: Inter-annotator agreement among the four hyperparameter optimization. LLMs used for sentiment analysis
3.6
As a further step to qualitatively assess the validity of sentiment votes, we plotted, for each tale, how the sentiment evolves during the narration (Figure 4). It is possible to observe, coherently with the patterns reported in the literature, how the evolution of the sentiment follows a trajectory of alternating states [22] which (with the exception of Frau Holle) culminates into a positive “happy end” [14].
3.5
Training
Our work relies on the XGBoost classifier [4]. The model preparation was conducted in two stages to ensure robust generalization across tales. In the first stage, we performed a grid search for the best hyper-parameters applying a partitioning into folds via a Stratified Group K-Fold (SGKF)4 approach, using macro-F1 as the selection criterion, with early stopping after 150 rounds, grouping by Tale to prevent data leakage between segments originating from the same instance. Stratification maintained the sentiment class distribution in each fold. The grid search was performed over a compact set of hyperparameters controlling model complexity, learning rate, and regularization: max depth, min child weight, eta, gamma, subsample, colsample bytree, lambda, alpha, scale pos weight. Table 2 reports the selected hyper-parameters values.
Evaluation Metrics
For each cross-validation fold, predictions were evaluated in terms of overall accuracy, balanced accuracy, precision, recall, and F1-score. Balanced accuracy, which is defined as the mean of per-class recall values, was used to mitigate bias toward more frequent sentiment categories. Precision, recall, and F1-scores were computed in both macro and weighted forms. Macro-averaged metrics assign equal weight to each class, reflecting the model’s ability to generalize across sentiments, 4 https://scikit-learn.org/stable/ while weighted averages account for class frequency modules/generated/sklearn.model_selection. for overall stability across folds. StratifiedGroupKFold.html 6
atic case was for neutral labels predicted as negative.
Table 2: Selected XGBoost hyper-parameters. Parameter
Selected Value
max depth min child weight eta gamma subsample colsample bytree lambda alpha scale pos weight
5 1.5 0.06 0.15 0.85 0.75 2.0 0.6 0.9
4.2
Baseline comparison
Given the seminal status of the research in sentiment analysis for sign language, we could not identify any suitable baseline for a direct comparison on sign language videos analysis. However, early work on sentiment prediction on text reported an overall Pearson correlation ρ = 0.65 [19, 1]. In our exIn the second stage, we trained the model using periments, after converting classification labels to a standard five-fold stratified cross-validation regression values, we computed a ρ = 0.529, indito obtain the final performance estimates. cating a lower performance with respect to the text analysis domain.
3.7
Feature selection 4.3
To discard confounding features and improve model accuracy, we sorted the prediction features according to their mean importance across fold. We then repeated training and test using only the top-N most important features (selecting groups of 16, 32, 64, 96, 128, and 160). The best performance was achieved using the first 96 features.
4
Results
4.1
Performance
Figure 7 reports the 30 most important features selected by the model training phase. In the following, we try to give some interpretation of their importance in the prediction of valence. Face smile and eyebrows. The contribution of the face in communicating emotions in SL has already been investigated, and thus the role of smile, eyebrows (raising or furrowing) or mouth corners (to smile or frown) is not surprising. Left elbow and shoulder distance from the camera. According to the mean depth coordinates of the left elbow and shoulder, the sentiment is more positive when the interpreter turns to her right. This might be explained by the fact that the interpreter often enacted negative sentences while impersonating an evil character, and thus rotating the torso to perform a role-shift. If true, this also suggests a tendency to place the novel’s antagonist always on the same side. The mean and standard deviation of the height of the right elbow have a positive correlation with valence, meaning that more positive sentences are characterized by higher and wider vertical movements of the right elbow (dominant). Hips vertical motion. There is a negative correlation between the valence and the standard deviation of the position, velocity, and acceleration of the hips on the vertical axis. This means that for negative sentences the interpreter is performing more variations of motion on her vertical axis. Indeed, it can be noticed from the video performances
Figure 5 shows the resulting main metrics across the five fold. Table 3 reports the numerical details. On average, our model reaches 63.1% balanced accuracy and 63.5% macro-F1. Figure 6 shows the confusion matrix for fold 1, the one with the highest accuracy, where it is visible that the most problem-
Sentiments count after majority vote 200 150 100 50 0
ti
mul
neg
e
ativ
tral
neu
Most important features
itive
pos
Figure 3: Distribution of sentiment among the 574 sentences of our dataset. 7
1-DHUDI
2-FrauHolle
3-DerWolf
4-Schneewittchen
Positive
Positive
Positive
Positive
Neutral
Neutral
Neutral
Neutral
Negative
Negative
Negative
Negative
S1 S6 S11 S17 S22 S27 S32 S37
S1 S11 S21 S33 S46 S56 S66 S76
S1 S10 S23 S33 S43 S52 S61 S71
Positive
Positive
Positive
Neutral
Neutral
Neutral
5-HaenselUndGretel
6-Dornroeschen
Negative
Negative S1 S14 S28 S41 S53 S66 S79 S92
S1 S21 S42 S62 S82 S106 S126
7-BremerStadtmusikanten
Negative S1 S9 S18 S27 S35 S44 S53 S61
S1 S11 S21 S31 S44 S55 S66 S76
Figure 4: For each fairy tale, the sentiment is plotted against the evolution of the plot. The dashed line shows the rolling mean on a 7-sample moving window. Table 3: Metrics computed on the 5 folds and their average. Acc.
Bal. Acc.
Prec.
Macro Recall
F1
Prec.
Weighted Recall F1
Neg.
Recall Neut.
Pos.
Pearson ρ
1 2 3 4 5
0.663 0.625 0.650 0.631 0.621
0.669 0.611 0.644 0.628 0.602
0.663 0.645 0.649 0.636 0.637
0.669 0.611 0.644 0.628 0.602
0.663 0.620 0.646 0.632 0.611
0.670 0.634 0.651 0.632 0.630
0.663 0.625 0.650 0.631 0.621
0.664 0.623 0.650 0.631 0.619
0.686 0.657 0.657 0.657 0.611
0.628 0.674 0.674 0.628 0.714
0.692 0.500 0.600 0.600 0.480
0.509 0.527 0.535 0.568 0.507
Mean
0.638
0.631
0.646
0.631
0.635
0.643
0.638
0.638
0.654
0.664
0.574
0.529
Fold
5-Fold Cross-Validation Metrics
1.0
96 features selected for the best performing model is available in appendix 8.1.
Score
0.8 0.6 0.4 0.2 0.0
5
Accuracy Balanced_Accuracy Macro_F1 Weighted_Precision Weighted_Recall 1
2
3
Fold
4
Limitations
One of the main limitations of this work is that no human annotators participated in the labeling task of the dataset. Although LLMs are recognized for their relatively high performance in sentiment analysis tasks, the contribution of human annotators would set a more reliable and validated ground truth. Furthermore, given that the main goal of the study is to predict the sentiment of SL videos, the most appropriate setup would include the annotation and agreement performed directly on the video material by native DGS annotators. As stated earlier, the signer is performing many role shifts, thus biasing the recognition of emotions with the recognition of role taking of a character potentially associate to positive or negative communications. For an unbiased extraction of the significant features involved in the sentiment predic-
5
Figure 5: Test metrics per folder.
how she is “jumping” on place when the context of the story gets troublesome. The mean distance between elbows has a positive correlation with valence. Suggesting that more positive utterances are performed by widening the movement of the arms. An exhaustive analysis and explanation of the correlations is beyond the scope of this work. For future linguistic analyses, the complete list of the 8
Negative
Confusion Matrix - Fold 1 0.69
0.17
0.14
within such long video sequences. To partially address this issue, we introduced only the “peak count” features, while other approaches may be available. An alternative approach would employ an continuous analysis on shorter moving time windows.
0.6
True Neutral
0.5 0.26
0.63
0.12
0.4
Second, for this study, we used only the frontal videos of the DGS-Fabeln-1 using the MediaPipe library, which is popular for its speed but less for its accuracy. Better body/face motion analysis could be achieved by using more accurate (although slower) systems (such as OpenPose: https://github.com/ CMU-Perceptual-Computing-Lab/openpose), or by simultaneously using the seven viewpoints available in the DGS-Fabeln-1 corpus, which would allow a triangulation for better landmarks and face blendshape estimation.
Positive
0.3 0.15
0.15
0.69
Negative
Neutral Predicted
Positive
0.2
Figure 6: Confusion matrix for test fold 1.
tion, torso rotation should be compensated for by A more comprehensive discussion about the readata normalization, or the stimuli material should sons why some of the features are so relevant for be selected to prevent role taking. the recognition of sentiment valence is beyond the scope of this work. However, we believe that our findings can help researchers in the lin6 Conclusions and Future presented guistic domain to better investigate sentiment expression in sign language. In general, we hope Work this is an inspiring work for future tight collaboraWe presented a method for the construction of tions between linguists and technical practitioners a corpus for sentiment analysis on German sign on using technology in the service of a better unlanguage fairy tales using LLMs, followed by the derstanding of sign languages, beyond the widely pipeline for training an explainable prediction spread construction of (uninterpretable) translamodel for the inference of sentiment valence from tion systems. videos. This is to our knowledge the first work performing a systematic analysis of both face and body contributions for sentiment expression using explainable models. The results show the ability to predict threelevel sentiment valence with a mean balanced accuracy of 63.1% and a macro-F1 score of 63.5%. About half of the most important features for the 7 Acknowledgements prediction belong to the body, which suggests the importance of joint face and body analysis for the prediction of valence from sign language utterances. This contribution is funded by the German MinOur research could be improved in several ways. istry for Education and Research (BMBF) through First, our analyses run on full video segments, last- the BIGEKO project (grant number 16SV9094), ing almost 10 seconds on average. However, sen- and by the Federal Ministry of Research, Technoltiment manifestation is often expressed by specific ogy and Space (BMFTR) through the RoGSiLT key signs, whose contribution might be absorbed project. 9
mouthSmileRight_mean mouthSmileLeft_mean browDownLeft_mean pose_LEFT_ELBOW_z_mean mouthUpperUpRight_mean pose_LEFT_HIP_y_velocity_std pose_RIGHT_HIP_y_acceleration_std pose_LEFT_SHOULDER_z_mean mouthRight_mean pose_RIGHT_HIP_y_velocity_std pose_LEFT_HIP_y_acceleration_std pose_RIGHT_ELBOW_y_std mouthDimpleRight_std mouthDimpleRight_mean jawRight_mean pose_RIGHT_HIP_y_std browDownRight_mean mouthLowerDownLeft_mean mouthSmileRight_std mouthUpperUpLeft_mean pose_RIGHT_ELBOW_y_mean browDownLeft_std pose_LEFT_HIP_y_std head_pitch_deg_mean pose_RIGHT_HIP_z_std browOuterUpLeft_mean dist_elbows_lr_avg mouthFrownLeft_std browOuterUpLeft_std left_hand_WRIST_y_mean
0.000
Top 30 Feature Importances
0.005
0.010
0.015
0.020
0.025
Importance mean and std
0.030
Figure 7: Top 30 features for the prediction of the sentiment labels. Mean and standard deviation among the five folds.
8
Optional Supplementary Materials: Appendices, Software, and Data
8.1
Appendices
Tables 4 and 5 list the features used to train the best model, sorted by importance. In the MediaPipe coordinate system: x is the horizontal axis, y is the vertical axis, and z increases with distance from the camera.
References [1] Buechel Sven and Hahn Udo. Emotion Analysis as a Regression Problem - Dimensional Models and Their Implications on Emotion Representation and Metrical Evaluation. In Frontiers in Artificial Intelligence and Applications. IOS Press, 2016. 10
[2] Necati Cihan Camgoz, Simon Hadfield, Oscar Koller, Hermann Ney, and Richard Bowden. Neural Sign Language Translation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), June 2018. [3] Santiago Castro. Fast Krippendorff: Fast computation of Krippendorff’s alpha agreement measure, 2017. Publication Title: GitHub repository. [4] Tianqi Chen and Carlos Guestrin. XGBoost: A Scalable Tree Boosting System. In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, pages 785–794, San Francisco California USA, August 2016. ACM. [5] Phoebe Chua, Cathy Mengying Fang, Takehiko Ohkawa, Raja Kushalnagar, Suranga Nanayakkara, and Pattie Maes. EmoSign: A Multimodal Dataset for Understanding Emotions in American Sign Language, 2025. Version Number: 1.
Table 4: List of the 96 prediction features sorted by importance. Part 1: 1-64. Feature mouthSmileRight mean mouthSmileLeft mean browDownLeft mean pose LEFT ELBOW z mean mouthUpperUpRight mean pose LEFT HIP y velocity std pose RIGHT HIP y acceleration std pose LEFT SHOULDER z mean mouthRight mean pose RIGHT HIP y velocity std pose LEFT HIP y acceleration std pose RIGHT ELBOW y std mouthDimpleRight std mouthDimpleRight mean jawRight mean pose RIGHT HIP y std browDownRight mean mouthLowerDownLeft mean mouthSmileRight std mouthUpperUpLeft mean pose RIGHT ELBOW y mean browDownLeft std pose LEFT HIP y std head pitch deg mean pose RIGHT HIP z std browOuterUpLeft mean dist elbows lr avg mouthFrownLeft std browOuterUpLeft std left hand WRIST y mean pose RIGHT HIP z mean eyeLookDownLeft mean eyeSquintLeft mean mouthShrugLower std pose RIGHT HIP y mean dist right wrist to nose avg head yaw deg mean pose LEFT SHOULDER z std mouthUpperUpRight std torso yaw mean mouthShrugLower mean pose LEFT SHOULDER x acceleration mean jawRight std browOuterUpRight mean cheekSquintRight mean mouthRight std mouthFrownRight mean right hand WRIST x acceleration std eyeLookDownRight mean pose NOSE z mean R WRIST accum dist avg pose LEFT HIP z std pose LEFT HIP z velocity mean mouthRollUpper mean pose LEFT ELBOW y std pose LEFT SHOULDER y velocity std right hand WRIST y mean mouthLowerDownLeft std browOuterUpRight std mouthLeft std browInnerUp std eyeBlinkLeft std dist right wrist to left shoulder avg mouthPressLeft peaks per s
Importance 0.02707 0.02354 0.01867 0.01661 0.01561 0.01530 0.01529 0.01475 0.01430 0.01419 0.01391 0.01343 0.01322 0.01305 0.01296 0.01278 0.01277 0.01266 0.01240 0.01211 0.01209 0.01190 0.01184 0.01173 0.01166 0.01161 0.01156 0.01133 0.01122 0.01106 0.01102 0.01089 0.01082 0.01073 0.01052 0.01044 0.01037 0.01030 0.01019 0.01015 0.01008 0.00991 0.00989 0.00987 0.00984 0.00984 0.00973 0.00970 0.00965 0.00963 0.00934 0.00932 0.00926 0.00926 0.00925 0.00917 0.00915 0.00914 0.00906 0.00897 0.00895 0.00888 0.00883 0.00879
11
. Table 5: List of the 96 prediction features sorted by importance. Part 2: 65-96. Feature mouthLeft peaks per s mouthFrownLeft mean pose RIGHT SHOULDER z mean eyeWideRight mean pose RIGHT HIP z velocity mean eyeBlinkRight std pose LEFT HIP y mean noseSneerRight std pose LEFT SHOULDER y acceleration mean mouthPressRight std eyeLookInLeft peaks per s L SHOULDER accum dist avg dist left wrist to left shoulder peaks per s pose NOSE x mean pose RIGHT SHOULDER z velocity std pose RIGHT HIP z acceleration std pose LEFT ELBOW y mean torso roll mean mouthPucker std pose RIGHT ELBOW y acceleration std mouthLowerDownRight peaks per s left arm angle mean pose LEFT ELBOW x mean right hand WRIST x velocity std pose RIGHT SHOULDER y mean left hand WRIST x acceleration std pose NOSE z velocity std mouthClose mean pose RIGHT SHOULDER x std right hand WRIST z velocity std pose LEFT ELBOW y peaks per s
Importance 0.00874 0.00867 0.00859 0.00858 0.00858 0.00857 0.00846 0.00846 0.00839 0.00839 0.00835 0.00825 0.00824 0.00819 0.00819 0.00816 0.00810 0.00807 0.00800 0.00790 0.00777 0.00775 0.00775 0.00760 0.00751 0.00751 0.00747 0.00734 0.00720 0.00702 0.00662
[6] David M. Eberhard, Gary F. Simons, and [13] Albert Mehrabian. Pleasure-arousalCharles D. Fennig. Ethnologue: Languages of dominance: A general framework for dethe World. SIL International, 28 edition, 2025. scribing and measuring individual differences in Temperament. Current Psychology, [7] Oliver Guhr, Anne-Kathrin Schumann, Frank 14(4):261–292, December 1996. Bahrmann, and Hans Joachim Böhme. Training a Broad-Coverage German Sentiment Clas- [14] Saif Mohammad. From Once Upon a Time sification Model for Dialog Systems. In to Happily Ever After: Tracking Emotions Nicoletta Calzolari, Frédéric Béchet, Philippe in Novels and Fairy Tales. In Kalliopi ZerBlache, Khalid Choukri, Christopher Cieri, vanou and Piroska Lendvai, editors, ProceedThierry Declerck, Sara Goggi, Hitoshi Isaings of the 5th ACL-HLT Workshop on Lanhara, Bente Maegaard, Joseph Mariani, guage Technology for Cultural Heritage, SoHélène Mazo, Asuncion Moreno, Jan Odijk, cial Sciences, and Humanities, pages 105–114, and Stelios Piperidis, editors, Proceedings of Portland, OR, USA, June 2011. Association the Twelfth Language Resources and Evaluafor Computational Linguistics. tion Conference, pages 1627–1632, Marseille, France, May 2020. European Language Re- [15] Fabrizio Nunnari, Eleftherios Avramidis, and Cristina España-Bonet. DGS-Fabeln-1, July sources Association. 2024. [8] Annika Herrmann, Sarah Schwarzenberg, Thomas Finkbeiner, Nina-Kristin Meister, [16] Fabrizio Nunnari, Eleftherios Avramidis, Cristina España-Bonet, Marco González, and Markus Steinbach. Expressing Emotions Anna Hennes, and Patrick Gebhard. DGSin Sign Languages (ExEmSiLa 2024), July Fabeln-1: A Multi-Angle Parallel Corpus of 2024. Fairy Tales between German Sign Language [9] Berenike Herrmann and Jana Lüdtke. A Fairy and German Text. In Proceedings of the Tale Gold Standard. Annotation and Analysis 2024 Joint International Conference on Comof Emotions in the Children’s and Household putational Linguistics, Language Resources Tales by the Brothers Grimm. 2023. Medium: and Evaluation (LREC-COLING 2024), pages HTML ,XML ,PDF Version Number: 1.0. 4847–4857, Torino, Italy, May 2024. ELRA and ICCL. [10] Jannis Klähn, Janos Borst-Graetz, and Manuel Burghardt. From dictionaries to LLMs [17] Osondu Oguike and Mpho Primus. Using – an evaluation of sentiment analysis techDeep Learning Models for Multimodal Senniques for German language data. Computatence Level Sentiment Analysis of Sign Lantional Humanities Research, 1:e4, 2025. guage. Forum for Linguistic Studies, April 2025. [11] Klaus Krippendorff. Content analysis: an introduction to its methodology. SAGE, Los [18] F. Pedregosa, G. Varoquaux, A. Gramfort, Angeles London New Delhi Singapore WashV. Michel, B. Thirion, O. Grisel, M. Blonington DC Melbourne, fourth edition edition, del, P. Prettenhofer, R. Weiss, V. Dubourg, 2019. J. Vanderplas, A. Passos, D. Cournapeau, M. Brucher, M. Perrot, and E. Duches[12] Camillo Lugaresi, Jiuqiang Tang, Hadon nay. Scikit-learn: Machine Learning in Nash, Chris McClanahan, Esha Uboweja, Python. Journal of Machine Learning ReMichael Hays, Fan Zhang, Chuo-Ling Chang, search, 12:2825–2830, 2011. Ming Guang Yong, Juhyun Lee, Wan-Teh Chang, Wei Hua, Manfred Georg, and [19] Daniel Preoţiuc-Pietro, H. Andrew Schwartz, Matthias Grundmann. MediaPipe: A Gregory Park, Johannes Eichstaedt, Margaret Framework for Building Perception Pipelines. Kern, Lyle Ungar, and Elisabeth Shulman. arXiv:1906.08172 [cs], June 2019. arXiv: Modelling Valence and Arousal in Facebook 1906.08172. posts. In Alexandra Balahur, Erik van der 12
Goot, Piek Vossen, and Andres Montoyo, editors, Proceedings of the 7th Workshop on Computational Approaches to Subjectivity, Sentiment and Social Media Analysis, pages 9–15, San Diego, California, June 2016. ACL. [20] Viktor Suter and Miriam Meckel. Using GPT4 for Text Analysis: Insights from English and German Language News Classification Tasks. ICWSM, US, June 2024. [21] Şeyma Takır, Barış Bilen, and Doğukan Arslan. Sentiment analysis in sign language. Signal, Image and Video Processing, 19(3):223, March 2025. [22] Ekaterina P. Volkova, Betty Mohler, Detmar Meurers, Dale Gerdemann, and Heinrich H. Bülthoff. Emotional Perception of Fairy Tales: Achieving Agreement in Emotion Annotation of Text. In Diana Inkpen and Carlo Strapparava, editors, Proceedings of the NAACL HLT 2010 Workshop on Computational Approaches to Analysis and Generation of Emotion in Text, pages 98–106, Los Angeles, CA, June 2010. Association for Computational Linguistics. [23] Weizhe Wang and Hongwu Yang. Towards Realizing Sign Language to Emotional Speech Conversion by Deep Learning. In 2021 12th International Symposium on Chinese Spoken Language Processing (ISCSLP), pages 1–5, Hong Kong, January 2021. IEEE. [24] Jiangtao Zhang, Qingshan Wang, and Qi Wang. U-Shaped Distribution Guided Sign Language Emotion Recognition With Semantic and Movement Features. IEEE Transactions on Affective Computing, 15(4):2180–2191, October 2024.
13