Man or machine: Evaluating Spelling Error Detection in Danish Newspaper Corpora

Eckhard Bick, Jonas Nygaard Blom, Marianne Rathje, Jørgen Schack

Research output: Chapter in Book/Report/Conference proceedingArticle in proceedingsResearchpeer-review

6 Downloads (Pure)

Abstract

This paper evaluates frequency and detection performance for both spelling and grammatical errors in a corpus of published Danish newspaper texts, comparing the results of three human proofreaders with those of an automatic system, DanProof. Adopting the error categorization scheme of the latter, we look at the accuracy of individual error types and their relative distribution over time, as well as the adequacy of suggested corrections. Finally, we discuss so-called artefact errors introduced by corpus processing, and the potential of DanProof as a corpus cleaning tool for identifying and correcting format conversion, OCR or other compilation errors. In the evaluation, with balanced F1-scores of 77.6 and 67.6 for 1999 texts and 2019 texts, respectively, DanProof achieved a higher recall and accuracy than the individual human annotators, and contributed the largest share of errors not detected by others (16.4% for 1999 and 23.6% for 2019). However, the human annotators had a significantly higher precision. Not counting artifacts, the overall error frequency in the corpus was low (~ 0.5%), and less than half in the newer texts compared to the older ones, a change that mostly concerned orthographical errors, with a correspondingly higher relative share of grammatical errors.

Original languageEnglish
Title of host publication3rd Annual Meeting of the ELRA-ISCA Special Interest Group on Under-Resourced Languages, SIGUL 2024 at LREC-COLING 2024 - Workshop Proceedings
EditorsMaite Melero, Sakriani Sakti, Claudia Soria
Number of pages8
Place of PublicationTorino
PublisherEuropean Language Resources Association (ELRA)
Publication date2024
Pages204-211
ISBN (Electronic)9782493814296
Publication statusPublished - 2024
Event3rd Annual Meeting of the ELRA-ISCA Special Interest Group on Under-Resourced Languages, SIGUL 2024 - Turin, Italy
Duration: 21. May 202422. May 2024

Conference

Conference3rd Annual Meeting of the ELRA-ISCA Special Interest Group on Under-Resourced Languages, SIGUL 2024
Country/TerritoryItaly
CityTurin
Period21/05/202422/05/2024

Bibliographical note

Publisher Copyright:
© 2024 ELRA Language Resource Association.

Keywords

  • Danish Newspaper corpora
  • Spell- and grammar checking
  • Spelling quality evaluation

Fingerprint

Dive into the research topics of 'Man or machine: Evaluating Spelling Error Detection in Danish Newspaper Corpora'. Together they form a unique fingerprint.
  • LREC-COLING

    Bick, E. (Participant)

    20. May 202422. May 2024

    Activity: Attending an eventConference organisation or participation

Cite this