Enhanced approach of multilabel learning for the Arabic aspect category detection of the hotel reviews

Asma Ameur*, Sana Hamdi, Sadok Ben Yahia

*Kontaktforfatter

Publikation: Bidrag til tidsskriftTidsskriftartikelForskningpeer review

15 Downloads (Pure)

Abstract

In many fields, like aspect category detection (ACD) in aspect-based sentiment analysis, it is necessary to label each instance with more than one label at the same time. This study tackles the multilabel classification problem in the ACD task for the Arabic language. For this purpose, we used Arabic hotel reviews from the SemEval-2016 dataset, comprising 13,113 annotated tuples provided for training (10,509) and testing (2,604). To extract valuable information, we first propose specific data preprocessing. Then, we suggest using the dynamic weighted loss function and a data augmentation method to fix the problem with this dataset's imbalance. Using two possible approaches, we develop new ways to find different categories of things in a review sentence. The first is based on classifier chains using machine learning models. The second is based on transfer learning using pretrained AraBERT fine-tuning for contextual representation. Our findings show that both approaches outperformed the related works for ACD on the Arabic SemEval-2016. Moreover, we observed that AraBERT fine-tuning performed much better and achieved a promising (Formula presented.) -score of (Formula presented.).

OriginalsprogEngelsk
Artikelnummere12609
TidsskriftComputational Intelligence
Vol/bind40
Udgave nummer1
Antal sider23
ISSN0824-7935
DOI
StatusUdgivet - feb. 2024

Bibliografisk note

Publisher Copyright:
© 2023 Wiley Periodicals LLC.

Fingeraftryk

Dyk ned i forskningsemnerne om 'Enhanced approach of multilabel learning for the Arabic aspect category detection of the hotel reviews'. Sammen danner de et unikt fingeraftryk.

Citationsformater