Abstract
Recurrent multiword units (RMUs) are central to language processing, yet their systematic identification and networked organization remain understudied. This study combines corpus-based and network-analytic methods to examine how RMUs contribute to the emergence of constructional schemas. Drawing on a 185-million-word corpus of Taiwan Mandarin, we pursue two aims. First, we propose a quantitative method for identifying cohesive RMUs based on word predictability in context. Second, we model RMUs as a network in which nodes represent RMUs and edges encode structural and semantic similarity, estimated with a state-of-the-art large language model. A comparison with a random sequence network confirms the non-random structure of the RMU network. Analysis of its topology reveals exemplar-based semantic groupings that support higher-level generalizations. These findings highlight RMUs as key building blocks in linguistic categorization, where subgroupings emerge through sequential lexical associations that underlie the formation of grammatical patterns and hierarchical structure.
| Original language | English |
|---|---|
| Pages (from-to) | 63-104 |
| Number of pages | 42 |
| Journal | Cognitive Linguistics |
| Volume | 37 |
| Issue number | 1 |
| DOIs | |
| Publication status | Published - 2026 Feb 1 |
Keywords
- constructional schema
- emergent grammar
- multiword units
- network analysis
- transitional probability
- usage-based grammar
ASJC Scopus subject areas
- Language and Linguistics
- Developmental and Educational Psychology
- Linguistics and Language
Fingerprint
Dive into the research topics of 'Recurrent multiword units as networks: sequentiality as basis for linguistic generalizations'. Together they form a unique fingerprint.Cite this
- APA
- Standard
- Harvard
- Vancouver
- Author
- BIBTEX
- RIS