Abstracts
Abstracts for oral presentations and posters at CMC-Corpora 2026.
Keynote abstracts are available on the Keynote Speakers page.
Oral Presentations
Establishing the Corpus of Danish Digitally Mediated Interaction (DanDIGI)
Torben Juel Jensen & Philip Diderichsen
University of Copenhagen, Denmark
The DanDIGI project establishes the first comprehensive, multimodal, cross-platform corpus of Danish digitally mediated interaction. Integrating data from five major social media platforms, the corpus combines large-scale public datasets with smaller, metadata-rich ethnographic collections. All material is encoded in TEI-conformant XML, preserving interactional structure, multimodal features, and platform-specific affordances. A dedicated processing pipeline comprises OCR, AI-assisted pseudonymization, structural markup, and linguistic annotation. Derived formats, including VRT and HTML-based context views, enable both quantitative and qualitative analyses within the same corpus infrastructure. DanDIGI provides a foundation for longitudinal studies, cross-modal comparisons, and data-driven research on contemporary Danish digital communication.
Back to programme
Rethinking Emoji Functions in Later Life: How Older Adults Interpret Emoji in Context
Lara Baert, Reinhild Vandekerckhove, Sarah Bernolet & Astrid De Wit
University of Antwerp, Belgium
Emoji are widely used in informal computer-mediated communication, yet their communicative functions remain difficult to operationalize, particularly for corpus annotation. Existing classifications are largely based on researcher interpretations and studies in younger populations. This study examines emoji function recognition among older social media users through a forced-choice classification task. The results show moderate overall interpretation of diverse emoji functions. While some functions were consistently confused with specific alternatives, suggesting broader interpretive domains, non-literal tone-modifying uses proved particularly difficult to identify. In these cases, responses did not converge on a single alternative category, suggesting a lack of shared interpretive patterns among older users. These findings suggest that some emoji functions, especially non-literal ones, may be less stably established in later-life communication than existing classifications assume. The study highlights the need for user-informed functional frameworks and has implications for the annotation of emoji use in CMC corpora.
Back to programme
Quantitative Approaches to German CMC Corpora
Andreas Witt & Marc Kupietz
Leibniz Institute for the German Language (IDS), Germany
This paper examines how quantitative linguistic measures can be adapted for corpora of computer-mediated communication (CMC) so that CMC data become comparable with spoken and written language. Such measures are well established in text linguistics, but they were developed for linear, monologic written texts and remain insufficiently operationalised for CMC. Drawing on recent work in text quality research and digital linguistics, we propose a set of measures covering lexical diversity, lexical density, segment length, part-of-speech distribution,and tense and mood distribution, complemented by interactional and multimodal metrics such as reply structures, hyperlinks, and emoji use. We sketch possible applications and discuss the open methodological questions.
Back to programme
Emojis in Swedish-Speakers’ WhatsApp Discussions: Hearts vs. Grins
Martti Mäkinen
Hanken School of Economics, Finland
This paper charts the use and variation of emojis in WhatsApp discussions by Swedish-speaking young adults in Finland, focusing on the most popular emoji elements, i.e. hearts, thumbs-ups, and grins. Variation in emoji use is observed with respect to the age, gender, bilinguality, and the home dialect of the data donors.
Back to programme
The Lexical Metaphor in Wine Speak: Semi-Automated, Manual Annotation in Expert and Non-Expert Wine Reviews
Léonieke Ariaans, Iris Hendrickx & Ilja Croijmans
Radboud University, The Netherlands
Tropes in wine discourse have increasingly attracted scholarly attention, yet systematic analyses of figurative strategies across expert and non-expert computer mediated communication remain limited. This paper examined how wine reviewers with different backgrounds (i.e. Vivino vs Wine Enthusiast) use multimodality concerning literal and figurative resources, focusing predominantly on lexical metaphors and their conceptual source domains that shape meaning. Findings showed that metaphorical language seemed to consistently rely on anthropomorphic projections (e.g. WINE IS A PERSON), whereas 40 percent concerned non-human conceptual metaphors. Besides the identification of differences in metaphorical display, wine experts (Wine Enthusiast reviewers) were found to use larger quantities of metaphor to describe wine compared to relative non-experts (Vivino users).
Back to programme
Personal Pronoun Declarations in CMC Networks: Tracing a Sociolinguistic Innovation in the Context of the Weak-Tie Hypothesis
Rahel Albicker
University of Eastern Finland, Finland
This study investigates the spread of personal pronoun declarations as a sociolinguistic innovation in large digital social media networks. Previous research has shown that linguistic innovations spread primarily via weak network ties, while strong ties inhibit language change. However, this difference has been shown to level out in networks larger than 50-60 nodes. Pronoun declarations as a sociolinguistic phenomenon have been studied with respect to social activity and online connectedness, but there are no studies that investigate the innovation in networks of different sizes and strengths. This study analyzes pronoun declarations in large digital social networks, considering both network strength and size to test whether the weak-tie hypothesis holds in the context of pronoun declarations on social media. The results suggest that contrary to previous findings, pronoun declarations seem to be more prevalent in strong-tie networks. The leveling effect of larger networks is nevertheless observed.
Back to programme
Graph Classification Explainability Meets Extremist Narrative Analysis
Dimitra Niaouri1, Meriem Hamdane2, Michele Linardi1 & Julien Longhi1
1 CY Cergy Paris Université, France; 2 ENSEA (École nationale supérieure de l'électronique et de ses applications), France
We propose a methodology to detect and explain extremist narratives in online text by combining Graph Neural Networks (GNNs) with explainability methods. We model text instances as graphs where nodes represent messages and edges capture semantic relationships, enabling the GNN to leverage both textual and structural information for
classification. To provide interpretability, we identify influential connections between messages, leveraging Deep Learning model explainability and Large Language Models to generate explanations about the in-group and out-group interactions. Preliminary experiments were conducted on a French Annotate Dataset, which includes online discourse
around three different topics, demonstrating that our approach not only enhances detection performance compared to extremist narrative agnostic classification but also sheds light on how extremist narratives propagate.
Back to programme
Leave and Let Die: Announcing a Departure from an Online Discussion
Lydia-Mai Ho-Dac1, Céline Poudat2, Ludovic Tanguy1, Cécile Fabre1 & Josette Rebeyrolle1
1 Université Toulouse–Jean Jaurès, CLLE, CNRS, France; 2 Université Côte d’Azur, BCL, CNRS, France
This study investigates how users announce their departure from online discussions and, more broadly, from the online communities that have developed within the Wikipedia project. Based on an automatic extraction of all posts in the English and French Wikipedia talk pages that are likely to contain such an announcement, we have conducted two analyses. The first aims to globally characterise the context in which these posts occur within the overall dynamics of the discussion: who, when, and what happens afterwards. The second expands on this analysis on a subsample with a manual annotation allowing a qualitative description of the type of departure announced, what is being left behind and the reasons why the user chooses to leave.
Back to programme
Computational Modeling of Register Variation in Positive Online Language
Maria Berger1, Yulia Clausen1 & Hannah Seemann2
1 Ruhr University Bochum, Germany; 2 University of Tübingen, Germany
While offensive language on social media has been extensively studied from both linguistic and computational perspectives, positive language has received considerably less attention. In this study, we examine three previously identified forms of positive online language: hope speech, candy speech, and empowering language. Building on Biber and Conrad’s (2019) framework, we conceptualize these forms of positive language as distinct registers defined by their communicative purposes. To examine the similarities and differences among these registers, we employ both traditional machine learning methods, such as SVM, and BERT-based language models. We extract features expected to be indicative of the studied registers such as sentiment and social media-specific features. This data-driven approach allows us to refine the theoretical boundaries of these closely related concepts and supports a more accurate automatic identification of distinct forms of positive language in online contexts.
Back to programme
Decoding Parameters as Methodological Choices: The Impact of Temperature on ASR-Derived Corpus Construction
Steven Coats
University of Oulu, Finland
Automatic speech recognition (ASR) is increasingly used for large-scale corpus construction, yet the impact
of decoding parameters on resulting datasets remains underexplored. We compare two operational decoding
configurations selected by fixed temperature settings (0.0 vs. 0.1) in a Singapore-English-adapted Whisper pipeline.
In a paired rerun of 1,322 podcast recordings (774 hours), sampling at temperature 0.1 produced 99,011 more
aligned words than beam search at temperature 0.0 (+1.02%; 95% paired-bootstrap interval: +0.89–1.16%).
Hand-checked passages show that these differences include both recovered audible material and acoustically
unsupported insertions. An additional sampling-only control at temperature 0.01 produced more output than at 0.1,
showing that output quantity was not monotonic with temperature. We argue that decoding configuration should be
treated as a methodological decision in corpus construction rather than a purely technical setting.
Back to programme
Sentence-Level Universal Dependencies (UD) Genres for CMC: A Social-Genre POS-Tagging Showcase
Egon W. Stemle
Eurac Research, Italy
Universal Dependencies (UD) is widely used as a multilingual infrastructure for morphosyntactic annotation and evaluation, and evaluation is typically reported at the level of whole treebank splits. For CMC-oriented questions, this is limiting, because many treebanks contain heterogeneous material while the resulting score is interpreted as a single evaluation condition. Building on earlier work that released sentence-level genre annotations for UD as a reproducible, versioned sidecar layer, this paper presents a first CMC-facing showcase of what such a new resource enables. We focus on the genre social as a practical proxy for CMC- and social-media-adjacent material and evaluate UDPipe models on full, social-only, and non-social test slices within the same underlying treebanks. The results do not support a plain "social is harder" conclusion but show that sentence-level genre labels expose evaluation contrasts that remain hidden under conventional whole-treebank scoring. The main contribution of the paper is therefore infrastructural and methodological: demonstrating that UD resources can be repurposed as reusable genre-sliced evaluation sets for CMC-oriented analysis.
Back to programme
Medium, Register, and the Individual as Sources of Linguistic Variation
Hannah Seemann1 & Tatjana Scheffler2
1 University of Tübingen, Germany; 2 Ruhr University Bochum, Germany
Present approaches of authorship analysis make use of stylometric methods that build on the assumption of an author's socio-demographic characteristics driving linguistic variation. In the context of computer-mediated communication, there are numerous additional factors that influence linguistic choice. We investigate two such factors relevant in the context of online communication: medium and register. Using features typically used as input for forensic analyses, word n-grams and part-of-speech, we show that the medium and in parts also register explain more variation between documents than the individual authors of the texts.
Back to programme
An Age-Cohort Approach for Analyzing X (Formerly Twitter) Data
Nobutane Hanayama & Motohiro Fukuoka
Shobi University, Japan
In the world of social networking services (SNS), many tweets or posts are considered to reflect the social conditions and events of their time. Accordingly, people may be interested in which conditions or events have a large impact and how long that impact persists. To examine the impact of such conditions or events, the numbers of “likes” and “retweets” or “reposts” may be considered useful indicators. However, it is difficult to determine the exact times at which the “like,” and “retweet,” or “repost” buttons were pressed, even when using the AI system Grok implemented in X (formerly Twitter). In this study, we regard an original tweet or post on X (formerly Twitter) as analogous to human survival and assume that it remains “alive” until the time of the last ”retweet” or ”repost”. We then propose a procedure for conducting survival analysis of tweets or posts using an age–cohort approach.
Back to programme
Regional and Minority Languages in CMC: Exploring Low German on Instagram
Frederike Schram
University of Turku, Finland
This paper examines the use of Low German in commercial Instagram posts to demonstrate how regional and minority languages in online spaces can be analysed using multimodal corpora. The described research project combines mixed-methods approaches of qualitative and quantitative analysis to investigate linguistic practices, associations with Low German, and strategies to construct regional identities. It draws on the concepts of enregisterment, postvernacularity, and commodification. The findings show that Low German is primarily used to index regionality and authenticity, while its communicative value is limited. The paper highlights the broader applicability of these approaches to other multilingual and RML contexts on social media.
Back to programme
The Flattening Effect: A Comparative Study of Sociopragmatic Erasure and Identity Loss in Translated Vietnamese-English Cyber-Discourse
Ming Ze Tang1, Phan Quynh Thu Nguyen2 & Thi Hong Nhung Nguyen2
1 Queen's University Belfast, United Kingdom; 2 University of Aberdeen, United Kingdom
This study investigates the "flattening effect" in translated Vietnamese-English code-mixed discourse, a phenomenon where semantic content is preserved while sociopragmatic nuances are erased. Utilizing a subset of the VietMix corpus, an "LLM-as-a-Judge" framework was employed to evaluate four dimensions: Structural Hybridity, Emotional Buffering, Status Signaling, and In-group Identity. While the model demonstrated moderate reliability in binary feature detection (inclusive κ = 0.416), its alignment with human consensus collapsed when quantifying erasure severity (expressive κ = 0.141). Qualitative analysis revealed that the model frequently over-interprets functional jargon while failing to recognize localized subcultural markers. These findings highlight a critical capability gap in current models, which prioritize global linguistic patterns over localized sociopragmatic functions. Consequently, future developments should move beyond semantic equivalence to incorporate localized cultural awareness within automated computer-mediated communication (CMC) analysis.
Back to programme
Corpus-Assisted Discourse Analysis on “Good English”: Attitudes and Ideologies on the Suomi24 Discussion Forum
Lea Meriläinen & Heli Paulasto
University of Eastern Finland, Finland
This study examines online discussions on English language skills and linguistic variation in the Suomi24 corpus covering the years 2021–2023. We apply corpus-assisted discourse analysis in order to identify evaluative discussions concerning English language skills and linguistic variation. In online discussion forum data, these themes are linked both to personal language attitudes and to broader, sociolinguistically prominent language ideologies. The way that English language skills are commented on and whose language skills are the focus of discussion reveal that Finns have an emotionally charged relationship with the English language. Our findings also indicate that online discussants’ awareness of linguistic variation is rather narrow, and good command of English is often equated with native-like proficiency.
Back to programme
The Self-Censorship Strategies of Crime Junkies Online: An Analysis of Algospeak Used by True Crime Content Creators on TikTok
Evelina Salo
Åbo Akademi University, Finland
Algospeak refers to linguistic strategies developed by social media users to avoid automated content moderation systems. As platforms intensify content governance, users increasingly adapt their linguistic practices to ensure visibility and avoid suppression, especially when discussing taboo or sensitive topics. This pilot study examines algospeak use within TikTok’s true crime community, a popular genre whose central topics of violence, trauma, and death, trigger heightened moderation on the platform. Drawing on a sample of eight videos from three creators, analyzed using Herring’s Computer-Mediated Discourse Analysis (CMDA), the study identifies key structural forms of algospeak and their pragmatic and social functions. Orthographic substitution, emoji replacement, lexical innovation, and silenced audio emerge as central strategies. The findings highlight how algorithmic constraints shape the linguistic landscape of TikTok, how creators maintain narrative cohesion while self-censoring, and how algospeak has become normalized within this genre-specific digital discourse community.
Back to programme
Beyond Agreement: Using LLM Reasoning to Evaluate Speech Acts and Politeness in Social Media
Susan Herring & Soyeon Lee
Indiana University, United States
This study evaluates the ability of a Large Language Model (LLM) to analyze discourse-pragmatic phenomena in social media threads by going beyond simple agreement with human-assigned annotations to incorporate the LLM’s reasoning. LLM reasoning is an emergent genre of AI-mediated communication that can be mined as both a source of insight and an object of analysis. Content analysis methods are employed to compare Gemini 2.5 Pro and human annotators along multiple dimensions with respect to their annotation of two Reddit threads for speech acts and politeness using Computer-Mediated Discourse Analysis (CMDA) methods. Taking its reasoning into account reveals Gemini to be performing these tasks at a higher level than strict human-LLM agreement metrics do; it shows the LLM making few actual errors of either code assignment or reasoning, except for over-coding politeness.
Back to programme
An Approach to LLM-Based Annotation of Reply Relations
Harald Lüngen & Laura Herzberg
Leibniz Institute for the German Language (IDS), Germany
This study looks at whether large language models (LLMs) can annotate reply relations in computer-mediated communication. We split the task into two steps, similar to anaphora resolution: target identification, that is finding the post that is being replied to, and relation classification, that is deciding what kind of reply it is. We use a German Wikipedia talk page corpus that already has a manually checked gold standard, with targets and relation types fully annotated, and we test two reasoning-oriented models, Claude Sonnet 4.6 and Qwen3-30B. Both models do a decent job at finding the right target. Classifying the type of reply is harder and separates the two models clearly: Claude handles all three tested types with balanced predictions, while Qwen assigns the large majority of instances to a single type of reply relation.
Back to programme
Posters
The Enhanced German Wikipedia Talk Pages Corpus: New Annotations for Emojis and Gender-Sensitive Language
Marc Kupietz, Nils Diewald & Laura Herzberg
Leibniz Institute for the German Language (IDS), Germany
This poster presents an enhanced version of the German Wikipedia Talk Pages Corpus, freely available at https://korap.ids-mannheim.de/instance/wiki, featuring improved tokenization and new annotation layers, relevant for CMC related research. One enhanced annotation layer is produced by the language-independent conllu-cmc tagger (https://github.com/KorAP/conllu-cmc), which tags all kinds of CMC-specific token types according the Empirist2015 specification, including ASCII emoticons, Unicode image emojis, and, since recently also wiki-template-based emoji representations. Unicode emojis are now also further annotated with group (e.g., smileys & emotion, people & body, symbols), subgroup, full name (including skin tone) and version from Unicode Technical Standard, queryable just like otherwise morphosyntactic features.
An entirely new annotation layer is provided by the conllu-gender tagger (https://github.com/KorAP/conllu-gender), which annotates German gender-inclusive spelling variants (e.g., asterisk form Schüler*innen, colon form Schüler:innen, the internal majuscule I) as well as emerging neo pronouns, supporting corpus-based research on gender-fair language change in German CMC.
Back to programme
Participatory CMC Corpus Building: A Museum-Based Approach to Data Donation and Public Engagement
Louis Cotgrove
Leibniz Institute for the German Language (IDS), Germany
The construction of corpora for computer-mediated communication (CMC) is increasingly constrained by issues of data access, platform restrictions, and ethical consent (Cal Varela et al. 2025). At the same time, research in citizen science and participatory approaches highlights the potential of involving non-experts in knowledge production (Hecker et al. 2018; Schaefer et al. 2021). This poster presents a system designed to collect CMC data through a public interface and demonstrates how visitor interactions can be translated into contributions to corpus linguistics.
We present an interactive terminal in an upcoming language museum setting that combines exploratory data analysis with a guided data donation workflow. Visitors can input excerpts of personal chat data and receive immediate visualizations, including frequency-based analyses and comparisons with existing resources such as DeReKo (Kupietz et al. 2023), MoCoDa2 (Beißwenger et al. 2020), and the NottDeuYTSch Corpus (Cotgrove 2023). Crucially, this exploratory mode is separated from the donation process: users first interact with their data locally and are then invited to opt into a consent-based donation via a dedicated collection interface.
The poster focuses on (i) the end-to-end workflow from user input to validated corpus contribution, (ii) the technical pipeline connecting the interactive terminal to the MoCoDa2 infrastructure, and (iii) design decisions that support informed consent and data quality. It also shows how the system is embedded in the museum environment to encourage participation while maintaining transparency about data use.
We argue that this approach enables participatory CMC corpus construction in CMC contexts, especially the collection of authentic data under transparent and ethically controlled conditions, whilst also fostering public engagement with linguistic research. The system serves as a model for future initiatives aiming to democratise data collection in linguistics and the digital humanities.
Back to programme
Cross-Cultural Computer-Mediated Culinary Communication: Affordances and Constraints of the Current Food Emoji
Lieke Verheijen, Ilja Croijmans & Samantha Revels
Radboud University, The Netherlands
Emoji are an important visual element of CMC and often claimed to be universal (Abel, 2019). Zooming in on food emoji (e.g., 🌭🧁🌶️), these are said to “connect people through shared experiences of culture-specific cuisine or universal foods” (George et al., 2023:186). But Unicode’s goal to “foster inclusion and participation in the digital world” (Unicode, 2025) conflicts with their criteria for new emoji, which prioritize universality and broad relevance across cultures, thus impeding representation of culture-specific foods. Is the food emoji inventory inclusive or biased towards certain cultures? This project investigates the affordances and constraints of the current food emoji, focusing on cultural diversity and exclusion, and uses a mixed-methods approach.
First, a content analysis reveals how food emoji in Unicode are distributed over different food cultures. Second, a large-scale Qualtrics survey explores cross-cultural differences in familiarity with, use of, and interpretations of food emoji, as well as differences in perceived representations of cultures by food emoji. For 30 random food emoji, participants are asked if they know it, use it, and what they think it means. Next, they describe a favourite or iconic national/regional/local food in words, try to visualize it in emoji, and rate the emoji inventory on its usefulness to do so. Respondents are from 18 countries (n ≈ 30 per country), based on Prolific availability, global cuisine reputation (TasteAtlas, 2025), and representation of different food cultures. An interaction between country and experienced communicative efficacy of food emoji is expected, where participants from some countries will have difficulties visualizing their food with emoji, suggesting that Unicode’s food emoji are not representative of the diversity of global food cultures. Still, users might devise creative solutions to deal with the limitations of the current emoji toolbox. Results, including statistical analyses, will be presented and implications will be discussed.
Back to programme
Komplett behindert?!?!? Token-Based Embeddings for Analysing Polysemy of the Word behindert (‘disabled’) in German CMC Corpora
Tim Feldmüller1 & Theresa Schweden2
1 Leibniz Institute for the German Language (IDS), Germany; 2 Johannes Gutenberg University Mainz, Germany
Linguistic conceptualisations of disability are subject to diachronic change. Recently, qualitative studies from conceptual history research have been supported by corpus-linguistic and distributional-semantic approaches (Author 2025; Author & Author accepted). Using German newspaper corpora, it was shown how, in the 2000s, concepts of disability as a medical problem were supplemented with disability as a dimension of diversity to be protected from discrimination.
An exploratory analysis of social media corpora reveals yet another dimension of German word behindert. Many user comments in the NottDeuYTSch YouTube corpus (Cotgrove 2023), for instance, show an offensive use of the word behindert ('disabled') by younger users (“are you completely disabled? It was a prank you spazz”, own translation).
With the present project, we aim to extend previous research in two respects: First, we will focus specifically on pejorative uses of disability-related lexis in German CMC corpora (Youtube, Wikipedia discussions, Twitter), contained in the German reference corpus (Kupietz et al. 2018). Second, we aim to methodologically advance a previous study (Author & Author accepted) based on classical word embeddings (Mikolov et al. 2013). To this end, we use token-based embeddings instead of type-based ones. We are thus able to generate one vector representation per token, allowing for a substantially better treatment of polysemy, as in the invective use of behindert on social media.
The poster will outline the study's design and initial findings, providing a basis for discussion. First analyses based on concordance samples (500 concordances from Wikipedia discussions, 500 concordances from Youtube) show that, using the XL-LEXEME architecture (Cassotti et al. 2023), different usages of the word can be clustered. For example, we find two large clusters that differentiate between a verbal use on the one hand (something/someone is obstructed by something/someone) and two attributive variants on the other (someone is disabled; something/someone is ~stupid).
Back to programme
cmc-tagger: A Tagger for Computer-Mediated Communication Corpora
Marc Kupietz, Nils Diewald & Harald Lüngen
Leibniz Institute for the German Language (IDS), Germany
This paper introduces cmc-tagger, an open-source CoNLL-U tagger for computer-mediated communication corpora. It annotates emojis, emoticons, hashtags, URLs, @-addresses, Wikipedia emoji templates, and action words in the XPOS column and, for Unicode emoji, adds FEATS metadata such as group, subgroup, qualification status, Unicode version, and normalized emoji name. Implemented as a reusable Unix-style filter, it can be used as a standalone tool or integrated into larger annotation chains with minimal overhead. The tool is already used for different CMC corpora provided via the corpus analysis platform KorAP.
Back to programme
Reddit Discourses on Generative AI Usage in Work
Minna Palo & Tiina Räisänen
University of Oulu, Finland
It is becoming clear that the defining disruptive technology of this generation is generative artificial intelligence (GenAI). It is changing the way people communicate and find, consume, and (re)produce information, which in turn is changing work practices and especially knowledge work. Sometimes called the fourth industrial revolution, this era of GenAI is causing polemic in online communities. Our research seeks to examine online discourses regarding work in the era of GenAI. To do so, we turn to a social media site where discussion and debate is encouraged, namely Reddit, the self-branded “heart of the internet”. In this presentation we look at selected posts and comments from subreddits dedicated to discussion on either work or artificial intelligence and use CDA 1) to identify and critically examine the discourses of GenAI and work and 2) to reflect on the global power structures at play in the discourses on work and GenAI. We take a critical discourse approach to our small corpus and discuss also the transformative effects of these discourses. Preliminary findings suggest discourse bubbles of utopian hopes and dystopian worries of a future without work and of concerns over the ethics and environmental toll of GenAI tools and their effects on the critical thinking skills of humans.
Back to programme
TEI SIG CMC: Infrastructure and Resources for the Encoding of CMC Corpora
Harald Lüngen1, Laura Herzberg1, Michael Beißwenger2 & Andreas Witt1
1 Leibniz Institute for the German Language (IDS), Germany; 2 University of Duisburg-Essen, Germany
This poster presents the activities, aims, and resources of the TEI SIG CMC, focusing on the recently established CMC module and Chapter 9 of the TEI Guidelines. It outlines encoding features, best practices, and ongoing corpus projects adopting the standard. Additionally, the poster highlights shared resources, including a wiki, and sample files and style sheets in a GitHub repository.
Back to programme
Representation of Minoritized Groups in Social Media Responses to UK Government Communication during COVID-19
Hanna Limatius1, Matt Gee2 & Robert Lawson2
1 University of Helsinki, Finland; 2 Birmingham City University, United Kingdom
Research on public crisis communication has highlighted the importance of studying publics’ reactions to crises (e.g., Coombs and Holladay, 2014). By reacting to authorities’ messaging during a crisis, citizens contribute to the shaping and interpretation of the crisis as it develops (Falco, 2022; Limatius and Koskela, 2024). In studying modern crises, social media discourse can provide important insights into what kind of language is used in such responses.
The poster presents a corpus-assisted critical discourse analysis (Baker, 2012) of UK-based social media communication during the COVID-19 pandemic. Specifically, we focus on the representation of minoritized ethnic groups in citizens’ responses to government social media accounts on Twitter (now ‘X’). By zooming in on tweets featuring the term BAME (‘Black, Asian and Minority Ethnic’), we address the following questions:
(1) How are minoritized ethnic groups referenced and discussed in COVID-19 tweets addressing UK government accounts?
(2) How is authorities’ communication regarding these groups evaluated by the citizens?
Our work draws on a larger corpus of tweets constructed as part of the TRAC:COVID research project (Kehoe et al. 2021). For this poster presentation, we focused on tweets responding to or mentioning seven official government accounts (ca. 12 million words). After establishing the range of terms related to ethnicity in the corpus, we decided to focus on the term BAME.
Our analysis of the tweets combines quantitative and qualitative perspectives. First, we look at the frequencies of words co-occurring with BAME to establish the contexts in which BAME groups are represented. Second, we conduct a critical discourse analysis (Fairclough, 2013; van Dijk, 1995) of the tweets. Our results show that most tweets criticize the UK government’s actions towards marginalized groups during COVID-19. However, some Twitter users’ responses to government communication also portray minorities as irresponsibly noncompliant, which contributes to racist and stereotyping discourses.
References
Paul Baker. 2012. Acceptable bias? Using corpus linguistics methods with critical discourse analysis. Critical Discourse Studies, 9(3), pages 247–256.
Timothy W. Coombs, and Sherry Jean Holladay. 2014. How publics react to crisis communication efforts: Comparing crisis response reactions across sub-arenas. Journal of Communication Management, 18(1), pages 40–57.
Norman Fairclough. 2013. Critical discourse analysis: Critical study of language (2nd ed.) London: Routledge.
Gaetano Falco. 2022. Communication of crisis or crisis of communication? The conflicting “voices” of the covid-19 pandemic across the world. Altre Modernità: Rivista di studi letterari e culturali. 28, pages 138–157.
Andrew Kehoe, Matt Gee, Robert Lawson, Mark McGlashan, and Tatiana Tkacukova. 2021. TRAC:COVID – Trust and Communication: A Coronavirus Online Visual Dashboard. Available online at https://traccovid.com.
Hanna Limatius and Merja Koskela. 2024. Voices in government crisis communication in the United States during the COVID-19 pandemic: A rhetorical arena perspective. VAKKI Publications, 16, pages 90–107.
Teun van Dijk. 1995. Aims of critical discourse analysis. Japanese Discourse, 1, pages 17–27.
Back to programme
Fine-Tuning Embeddings for Register Classification Reveals Functional Subtypes within Social Media Registers
Erik Henriksson, Tuomas Lundberg & Veronika Laippala
University of Turku, Finland
Registers—language varieties associated with recurring linguistic patterns in specific contexts (Biber & Conrad 2019)—are often described in broad categories, though texts within a single register frequently exhibit systematic internal variation. These differences may reflect distinct subregisters, topical groupings, or gradual shifts along register continua (Biber & Egbert 2023), but identifying such patterns typically requires labor-intensive, corpus-specific analysis. This paper presents a computational approach to uncovering fine-grained structure within web-based registers, focusing on social media. To this end, we combine supervised register classification with unsupervised clustering. We investigate the types of social media patterns that emerge and their cross-linguistic consistency.
We fine-tune multilingual XLM-R models on a 25-class register taxonomy (Henriksson et al. 2024) to classify one million documents per language from English, Finnish, and Swedish web corpora drawn from the HPLT dataset (Burchell et al. 2024). This produces large register-specific subsets (e.g., ~490,000 Narrative Blogs, ~79,000 Interactive Discussions), with 5–15% of documents assigned multiple registers depending on language. Clustering the resulting embeddings reveals both topical and functional subtypes. For instance, Finnish Narrative Blog + Opinion texts divide into consumption- versus culture-oriented clusters, while sports-related clusters emerge in English Interactive Discussion. Functionally, clusters capture gradients of interactivity: across all languages, Narrative Blogs separate by degree of reader engagement, with comment-enabled blogs forming distinct hybrid clusters.
We validate these findings using UMAP visualization and keyness analyses, showing that clusters correspond to distinctive linguistic features. SVM classifiers predict cluster membership with 95–99% accuracy, indicating robust, learnable patterns. In contrast, clustering baseline pretrained embeddings yields less meaningful groupings driven by formatting and boilerplate. Fine-tuning thus reshapes embedding space to better capture nuanced register variation.
References
Biber, Douglas, and Susan Conrad. 2019. Register, Genre, and Style. Cambridge University Press.
Biber, D., Egbert, J. 2023. “What is a register? Accounting for linguistic and situational variation within – and outside of – textual varieties”. Register Studies 5:1, 1-22. https://doi.org/10.1075/rs.00004.bib
Burchell, L. et al. 2025. “An expanded massive multilingual dataset for high-performance language technologies (HPLT)”. https://doi.org/10.48550/arXiv.2503.10267.
Conneau, A., Khandelwal, K., Goyal, N., Chaudhary, V., Wenzek, G., Guzmán, F., Grave, E., Ott, M., Zettlemoyer, L., Stoyanov, V. 2019. “Unsupervised cross-lingual representation learning at scale”. https://doi.org/10.48550/arXiv.1911.02116.
Henriksson, E., Myntti, A., Hellström, S., Eskelinen, A., Erten-Johansson, S., Laippala, V. 2024. “Automatic register identification for the open web using multilingual deep learning.” https://doi.org/10.48550/arXiv.2406.19892.
McInnes, L., Healy, J. and Melville, J. 2018. “Umap: Uniform manifold approximation and projection for dimension reduction”. https://doi.org/10.48550/arXiv.1802.03426.
Back to programme
Design and Annotation of a Multilayer CMC Corpus for the Analysis of Spanish Conditional Constructions
Laura Aubry
University of Neuchâtel, Switzerland
Conditional constructions in Spanish constitute a complex linguistic category, exhibiting a wide range of semantic and pragmatic functions as well as a close syntactic interdependence between their main components. Their form, distribution and interpretation are also known to vary depending on communicative contexts (e.g. Dancygier and Sweetser, 2006). For example, studies have shown that the form of conditional constructions may vary across spoken/written registers, with spoken language often favoring simplified or non-canonical patterns (e.g. NGLE; Montolío, 1999; Santana Marrero, 2003). As a type of written digital communication displaying features associating with both conceptual orality and conceptual literacy, computer-mediated communication (CMC) provides a particularly relevant setting for observing such variation. However, despite this, the use of conditional constructions in CMC remains underexplored, partly due to the lack of adequately annotated resources.
This poster presents a proposal for the design and annotation of a multilayer annotated corpus of Spanish CMC, specifically developed to investigate conditional constructions. The annotation workflow combines normalization of non-standard forms where necessary with automatic lemmatization and morphosyntactic tagging, and includes manual annotation of conditional occurrences, implemented in a TEI-based XML schema. The annotation scheme integrates morphosyntactic, semantic (e.g. factual, hypothetical, counterfactual), and pragmatic (e.g. speech acts, politeness) levels. In addition, contextual metadata (e.g. degree of familiarity between participants, communication channel, public vs. private setting) are encoded where relevant, enabling corpus partitioning and quantitative analyses based on parameters related to communicative proximity (Koch and Oesterreicher, 2007).
The proposal discusses annotation guidelines and methodological challenges involved in the annotation of conditional constructions in CMC, and explores the potential of this approach for the systematic analysis of their form-function relationships in CMC discourse.
Back to programme
Setting Experiences against Expectations in Personal Narratives on Talking about Weight in Healthcare Encounters
Maarit Siromaa & Mirka Rauniomaa
University of Oulu, Finland
Drawing on personal narratives and an understanding of their being constructed in interaction to serve interactional purposes, the study examines the kinds of accounts patients provide of their discussions with healthcare providers about the patient's weight. The data consist of written narratives shared on closed discussion forums in social media which deal with people’s personal experiences and anecdotes of talking about weight in healthcare encounters. The data are analyzed with the help of qualitative methods based on research on narratives, language and social interaction. The study explores how the teller constructs their narrative in systematic ways to highlight key moments of healthcare encounters as bearing particular relevance for them and to give meaning to different aspects of weight talk in such encounters. Furthermore, the study explores how the teller positions themselves with regard to weight talk, their co-participants and possible others in healthcare encounters. The study focuses on two storytelling practices used by teller-patients: 1) narrative prefaces that align with or deviate from, often negative, expectations of healthcare encounters and 2) stretches of reported speech assigned by teller-patients to healthcare providers that display both patient and provider perspectives.
Back to programme
Integrating CMC into a Transversal Metadata Scheme: The Open French Corpus Case
Céline Poudat1, Christophe Parisse2, Mathilde Guernut3 & Achille Falaise4
1 Université Côte d’Azur, CNRS, BCL, France; 2 Université Paris Nanterre, CNRS, MoDyCo, France; 3 CORLI Consortium, Huma-Num, France; 4 Université Paris Cité, CNRS, LLF, France
For the past four years, the CORLI consortium has been developing the Open French Corpus (OFC) project (Parisse et al., 2023), which brings together existing French-language corpora within a standardized framework, both in terms of formats and metadata. The project aims to centralize these corpora in a shared, accessible environment on the Ortolang platform, supported by dedicated tools for exploration (Poudat et al., 2025). A controlled vocabulary for its metadata is being developed in parallel on the Opentheso platform.
CMC corpora represent a substantial share of the resources currently being integrated into OFC. A significant part comes from the CoMeRe repository (Chanier et al., 2014), which gathers 13 corpora across a wide range of CMC genres: SMS, tweets, wiki discussions, weblogs, text chats, and multimodal environments including 3D virtual worlds. This contribution reports on the metadata adjustments we had to make when integrating these diverse resources into OFC's transversal metadata scheme.
A transversal metadata framework necessarily loses some genre-specific detail, in exchange for broader interoperability. We illustrate this tension through concrete examples drawn from the integration of CoMeRe into OFC, focusing on two salient issues: (i) the typology of CMC corpora, where genre, platform and mode interact in ways that a single-axis classification hardly captures, and (ii) the communicative sphere, whose binary private/public distinction does not accommodate institutional or learning contexts. We further relate these issues to the recommendations of the TEI P5 CMC module (chapter 9), which provides an external reference.
Rather than offering a definitive answer, this poster raises questions intended to contribute to ongoing reflection on metadata design for heterogeneous CMC corpus infrastructures.
Back to programme
A Curation Strategy for the CLARIN Resource Family for Computer-Mediated Communication
Alexander König1, Egon W. Stemle2, Lionel Nicolas2, Céline Poudat3 & Steven Coats4
1 CLARIN ERIC, The Netherlands; 2 Eurac Research, Italy; 3 Université Côte d’Azur, France; 4 University of Oulu, Finland
The CLARIN Resource Families (CRF) are curated lists of high-quality resources for various domains identified as particularly relevant to the CLARIN community, one of which is the CRF for Computer-Mediated Communication (CMC).
Following the CRF coordinators' decision to request support from the CLARIN Knowledge Centres to help with curating these ever-growing lists, the CLARIN Knowledge-centre for Computer-Mediated Communication and Social Media Corpora has been asked to “adopt” the CMC CRF and thus take care of the task of adding new relevant entries, removing entries that no longer meet the inclusion criteria or are deemed not relevant enough anymore and updating existing entries if needed.
We introduced the core concept at last year’s CMC conference as a featured presentation and took the occasion to invite participants to join a taskforce tackling this effort. With this poster, we aim at presenting our first version of the curation guidelines.
We will detail the rationales behind the procedures devised and the criteria adopted on adding, updating and removing entries to the CRF. We will also take the occasion to describe the technical and organisational process as a whole and showcase the CMC-CRF itself. Finally we will argue in favor of not only showcasing popular datasets but also explicitly trying to include datasets covering unusual languages, unusual platforms or novel approaches.
Back to programme
“Yet Another Dumbass 🙂”: Understanding 🙂 within the Vietnamese Online Community
Penelope Nguyen
University of Sheffield, United Kingdom
How an emoji is used and interpreted may differ from culture to culture (Kejriwal et al., 2021). According to Emojipedia, the slightly smiling face emoji is ambiguous in that it can either convey positive emotions or sarcasm (🙂 Slightly Smiling Face Emoji, n.d.). This study sheds light on how 🙂 is used within the Vietnamese online community: how this emoji is used in posts and comments, whether the positive or negative connotation is more prevalent, and how the connotation is conveyed. To answer these questions, a two-pronged study using both corpus and interview data was conducted, First, a corpus was built of posts and comments on a Facebook confession page, known for a high density of impoliteness. Subsequently, a subcorpus (489,000 tokens) with only posts and comments containing 🙂 was created for both quantitative and qualitative analyses. Results based on raw frequency counts and mutual information (MI) scores show that (1) among all emojis present in the corpus, 🙂 was the third most popular, and (2) if considered a “node word,” it was most strongly associated with sentence-ending emotive particles and the evaluative adjective sai ‘wrong’. A qualitative analysis of 200 first comments in the 🙂 subcorpus, as well as a comparison with its upside-down version 🙃, offers insights into Vietnamese Facebook users’ habits of utilizing 🙂. Most posts and comments with 🙂 had negative connotations, specifically impoliteness, with some rapport mismanagement (Spencer-Oatey, 2000). Within Culpeper’s (2011, pp. 155–156) impoliteness framework, the emoji 🙂 creates internal mismatches to express implicational impoliteness. Simultaneously, an interview with 51 native Vietnamese active Facebook users reveals that the interviewees generally view 🙂 as a sarcastic and insincere smile, which is used to reflect negative attitudes. This aligns with and complements findings from the corpus analysis.
Back to programme
Performing Moral Authority: How ‘should be ashamed’ Circulates across Contrasting Affordances of X and Bluesky
Daria Dayter & Svetlana Filon
Tampere University, Finland
This study examines how moral authority is performed through the phrase “should be ashamed” across two contrasting platform ecologies: X, with its algorithmic amplification, and Bluesky, which currently lacks such mechanisms. Using a dataset of 400 posts (200 from each platform) collected via algorithmic ethnography-guided searches (Seaver 2017) conducted from newly created, unpersonalised accounts, the analysis identifies the actors positioned as subjects and objects of moral judgment. Drawing on Garcés-Conejos Blitvich’s (2022) work for understanding online public shaming, the study explores how denunciatory practices unfold differently depending on platform affordances. Haugh’s (2024) work on (im)politeness in online public forums informs the analysis of stance-taking, facework, and the negotiation of moral authority. Findings are expected to shed light on how the circulation, uptake, and interactional consequences of “should be ashamed” are shaped by user intent, but also by the technological architectures that structure visibility and engagement across networked publics.
Back to programme
CMC Knowledge Centre workshop: Metadata practices for CMC corpora
Alexander König1 & Egon W. Stemle2
1 CLARIN ERIC, The Netherlands; 2 Eurac Research, Italy
It is nowadays consensually accepted that standardised metadata contribute significantly to the FAIRification of a research field
and accordingly make the reuse of existing data in new research projects easier.
With some members of the CKCMC having worked on a similar task for the Learner Corpus Research community, we plan to use this
experience and take up the task of developing a CMC metadata schema
in a dedicated taskforce which will commence work in the autumn of this year.
This roundtable is meant to gather input from the community and recruit additional taskforce members.
We will outline our first thoughts on the matter and then discuss with you which metadata
fields should be part of a core CMC profile and which are more peripheral.
Back to programme