Improving Short Text Classification Through Global Augmentation Methods

Vukosi Marivate; Tshephisho Sefara

doi:10.1007/978-3-030-57321-8_21

Conference Papers Year : 2020

Improving Short Text Classification Through Global Augmentation Methods

(1, 2) , (3)

1
2
3

Vukosi Marivate

Function : Author
PersonId : 1098364

UP - University of Pretoria [South Africa]

CSRI - Council for Scientific and Industrial Research [South Africa]

Tshephisho Sefara

Function : Author
PersonId : 1115839

CSIR - Council for Scientific and Industrial Research [Pretoria]

Abstract

We study the effect of different approaches to text augmentation. To do this we use three datasets that include social media and formal text in the form of news articles. Our goal is to provide insights for practitioners and researchers on making choices for augmentation for classification use cases. We observe that Word2Vec-based augmentation is a viable option when one does not have access to a formal synonym model (like WordNet-based augmentation). The use of mixup further improves performance of all text based augmentations and reduces the effects of overfitting on a tested deep learning model. Round-trip translation with a translation service proves to be harder to use due to cost and as such is less accessible for both normal and low resource use-cases.

Keywords

Domains

Fichier principal

497121_1_En_21_Chapter.pdf (532.55 Ko)

Origin	Files produced by the author(s)
licence	CC BY 4.0 - Attribution

Connect in order to contact the contributor

https://inria.hal.science/hal-03414750

Submitted on : Thursday, November 4, 2021-3:58:34 PM

Last modification on : Wednesday, November 13, 2024-4:58:02 PM

Long-term archiving on : Saturday, February 5, 2022-7:10:49 PM

Dates and versions

hal-03414750 , version 1 (04-11-2021)

Licence

CC BY 4.0 - Attribution

Identifiers

HAL Id : hal-03414750 , version 1
DOI : 10.1007/978-3-030-57321-8_21

Cite

Vukosi Marivate, Tshephisho Sefara. Improving Short Text Classification Through Global Augmentation Methods. 4th International Cross-Domain Conference for Machine Learning and Knowledge Extraction (CD-MAKE), Aug 2020, Dublin, Ireland. pp.385-399, ⟨10.1007/978-3-030-57321-8_21⟩. ⟨hal-03414750⟩

Improving Short Text Classification Through Global Augmentation Methods

Abstract

Keywords

Domains

Dates and versions

Licence

Identifiers

Cite

Export

Collections

Altmetric

Share