Skip to main navigation Skip to search Skip to main content

Sustainable Development Goals in Corporate Reporting: Investigating the Incorporation of SDGs with Multi-label Text Classification

Per Axel HÃ¥kansson & Wiktoria Lazarczyk

Student thesis: Master thesis

Abstract

With the introduction of the 2015 United Nations Sustainable Development Goals, companies aimed to enhance their corporate social responsibility initiatives, generating a growing volume of reports and selfdisclosed information. The abundance of textual data resulting from this, creates a need for the automatic identification of SDG incorporation in CSR reporting. This thesis examines the question: "To what extent is the effectiveness of using Natural Language Processing (NLP) in identifying the incorporation of SDGs in CSR strategies through Corporate reporting?" To achieve this objective, we establish a project scope that involves using authorized SDG documents sourced from the United Nations and official corporate sustainability reports. This involves multi-label classification following a two-step approach: reports annotation and annotated reports classification. Reports annotation entails the combination of machine and deep learning models to investigate their knowledge transferring capabilities for reports data. The first dataset consists of SDG data obtained from UN documents, while the second dataset combines SDGs with samples of extracted reports. The goal is to acknowledge whether incorporating samples of the target domain to the training data improves the knowledge transfer. The results of our study demonstrate that the dataset comprising both SDGs and report sample data exhibits superior performance in both machine and deep learning compared to the SDG-only dataset. The RoBERTa model was found to have the highest accuracy (0.66) and F1 micro (0.67) among the chosen models. These results suggest that transfer-based models, specifically RoBERTa, outperform the other utilized models. Based on the learnings from the RoBERTa model, the next step involved the annotated report’s multi-label classification and examining its performance when no distinction between the source and target domain exists, employing BERT and RoBERTa. The RoBERTa model performs better in this endeavor than the pre-existing BERT model at all levels, achieving an accuracy of 0.91 and an F1 micro score of 0.92. These results suggest that the model performs well in predicting class membership. However, the F1 micro score of 0.42 indicates that the model's performance is limited in distinguishing certain classes, leading to underrepresented predictions. Therefore, for future research, the need for more robust data augmentation is highlighted. The study confirms the relevance and explorational power of transformer- based models in multi-label SDG classification and uncovers the importance of high quality and sizeable data.

EducationsMSc in Business Administration and Data Science, (Graduate Programme) Final Thesis
LanguageEnglish
Publication date15 May 2023
Number of pages130
SupervisorsNiels Buus Lassen