Abstract
This thesis probes whether three prevalent attention variants, additive, dot-product and multi-head self-attention, boost recognition accuracy and yield interpretable attention maps when grafted onto the same pose-based encoder for isolated German Sign Language (DGS) recognition. Using 811 frontal-view videos from the multi-signer, heavily imbalanced MEINE-DGS corpus, skeletal key-points extracted with OpenPose are spatially normalized and fed to four comparably sized encoders: a bidirectional LSTM baseline, the two LSTM-attention variants, and a compact two-layer Transformer. Models are evaluated on Top-1/Top-5 accuracy, macro-F1, and several attention-quality metrics. Results show that attention mechanisms do not provide clear advantages. The Transformer edges the baseline by only 0.1 percent points in Top-1 accuracy (16.0 % vs 15.9 %) while trailing it in Top-5 and macro-F1; both LSTM-attention models underperform across all metrics. Attention maps remain diffuse or misaligned. The study therefore concludes that, under extreme class imbalance and noisy key-point inputs, unguided frame-level attention neither enhances performance nor improves interpretability. Future work should explore richer representations, structured or region-aware attention, and stronger pre-training to unlock attention’s potential in sign-language recognition.
| Educations | MSc in Business Administration and Data Science, (Graduate Programme) Final Thesis |
|---|---|
| Language | English |
| Publication date | 15 May 2025 |
| Number of pages | 90 |
| Supervisors | Niels Buus Lassen |