What Small Corpora Can Still Tell Us

Authors

  • Leila Haddad Institute for Language Data Author

Keywords:

low-resource NLP, corpus linguistics, evaluation

Abstract

Large-scale benchmarks often hide the interpretive value of carefully constructed small corpora. We present a workflow for combining close linguistic annotation with lightweight statistical analysis and show how modest datasets can expose systematic variation that disappears in aggregated evaluation.

Published

2025-10-15