Papers
Topics
Authors
Recent
Search
2000 character limit reached

DinG Corpus: French Dialogue & AMR Resource

Updated 9 July 2026
  • DinG Corpus is a freely shareable collection of long, multi-party French dialogues captured during board game sessions.
  • It employs detailed manual transcription with explicit time codes, overlap handling, and a robust question categorization scheme.
  • The ding-01 extension adds AMR annotations, transforming the corpus into a benchmark for semantic parsing in spontaneous dialogue.

Searching arXiv for the DinG corpus and related French dialogue/AMR resources. DinG, short for “Dialogues in Games,” is a freely shareable corpus of long, spontaneous, real-life multi-party dialogues in French, recorded while groups play the board game Catan. It was created to provide a high-quality resource for studying dialogue phenomena in French, particularly questions and turn-taking in natural conversation, while mitigating privacy concerns by centering interaction on gameplay rather than personal disclosure (Boritchev et al., 2022). A subsequent extension, ding-01, adds Abstract Meaning Representation (AMR) annotations to a subset of DinG, turning the corpus into a semantic resource for spontaneous French dialogue and a benchmark for speech-like semantic parsing (Kang et al., 18 Aug 2025).

1. Origin, scope, and research objectives

DinG was designed for “long, oral, spontaneous, multi-party dialogues in French,” a setting for which publicly shareable resources are comparatively scarce. Its immediate motivation was to support the study of discourse phenomena, dialog acts, and interaction dynamics in French, with particular emphasis on questions in dialogue and on comparisons with discourse-level modeling projects such as STAC, which concerns Catan dialogues in a different modality and language (Boritchev et al., 2022).

The choice of Catan is methodologically central. As a resource-trading game, it elicits negotiation, requests, offers, confirmations, and turn management under naturally recurring task constraints. The game-centered setup also reduces personal talk, which in ordinary dialogue corpora often blocks open distribution. The published transcriptions therefore preserve spontaneous, real-life interaction while remaining compatible with broad reuse under a free license (Boritchev et al., 2022).

The corpus is notable for combining several properties within a single French resource: it is oral rather than text-based, multi-party rather than dyadic, long rather than composed of isolated utterances, and openly shareable rather than restricted by sensitive content. This combination makes DinG relevant to conversational analysis, discourse studies, question classification, ASR post-processing, and dialogue-system research (Boritchev et al., 2022).

2. Data collection, participants, and recording conditions

The recordings were made during university game nights in the same room as other games, with the explicit goal of keeping participants relaxed and natural. A non-player observer explained the rules, recruited volunteers, and supervised each session. Before each recording, rules were explained and a rule book was provided so players could play autonomously (Boritchev et al., 2022).

The participant pool comprises 33 speakers, including 12 women and 21 men. Each speaker appears once. The average age is approximately 25 years, all speakers are native French speakers, and all but 3 had a master’s degree or higher. No specific French regional variants are reported (Boritchev et al., 2022).

The resource contains 10 recorded games. A typical Catan game involves 3–4 players, and most recordings—specifically all but numbers 4, 5, and 6—were split into two parts because of a food break. The total duration is 701.8 minutes, approximately 11.7 hours, with per-game lengths ranging from 39.92 to 104.33 minutes and an average of approximately 70 minutes (Boritchev et al., 2022).

The audio was captured in a noisy, real-life environment that included dice rolls, table noise, and conversations from other games. Audacity was used for preprocessing, including reducing peaks such as dice rolls, uniform amplification, and background noise reduction. After evaluating several transcription tools, the project selected ELAN for segmentation, transcription, time-code alignment, and overlap visualization (Boritchev et al., 2022).

3. Transcription architecture and corpus structure

DinG is a corpus of manual transcriptions with explicit time alignment and overlap handling. The transcription process involved 6 trained transcribers—5 NLP students and 1 subtitling expert—who were trained on a common 5-minute excerpt. Transcribing a 1.5-hour recording took approximately 30 hours (Boritchev et al., 2022).

Segmentation is manual and organized into “speech turns,” each defined as a coherent linguistic unit. These units may be pseudo-sentences, onomatopoeia, noise, or combinations thereof. Overlaps are explicitly represented and disambiguated in ELAN. Each segment has start and end time codes, and gaps between non-overlapping segments are automatically computed and shown in brackets in the exported text files (Boritchev et al., 2022).

The transcription conventions encode interrogative marks explicitly, annotate pauses with durations, and include laughter, onomatopoeias, and gameplay-adapted noise tags such as [dice] and [tokens]. Disfluencies and syntax received automatic annotations using SLODiM. This makes the corpus suitable not only for discourse analysis but also for research on transcription, spoken-language phenomena, and the interaction between manual and automatic annotation layers (Boritchev et al., 2022).

Speaker identification is anonymized. Speakers are labeled by token color—R, W, Y, and B—with “O” reserved for outside noises or outside speakers. Names are removed, and if a name is mentioned in dialogue it is replaced with the corresponding color in uppercase. Only transcriptions are published; audio is not currently released (Boritchev et al., 2022).

A concise summary of the corpus scale is as follows.

Property DinG value Source
Recorded games 10 (Boritchev et al., 2022)
Total duration 701.8 minutes (Boritchev et al., 2022)
Total turns 23,575 (Boritchev et al., 2022)
Total questions 2,528 (Boritchev et al., 2022)
Average turns per minute 33.59 (Boritchev et al., 2022)
Average questions per minute 3.60 (Boritchev et al., 2022)
Questions among turns 10.72% (Boritchev et al., 2022)

The corpus is described as fairly homogeneous across games, especially in the proportion of questions among turns, with a coefficient of variation of 20% for the percentage of questions among turns. By contrast, length, number of questions, and turn rate vary more substantially across games, reflecting differences in game tempo and interactional density (Boritchev et al., 2022).

4. Question-centered annotation and interactional analysis

A major research focus of DinG is the study of questions in dialogue. Questions were automatically retrieved via question marks and extracted together with two turns of left and right context. The annotation scheme, adapted from Cruz Blandon et al. (2019), distinguishes six categories: YN (yes/no-question), WH (wh-question), DQ (disjunctive question), CS (completion suggestion), PQ (phatic question), and N/A (non-assigned) (Boritchev et al., 2022).

Automatic pre-annotation followed four ordered rules: if the next utterance starts with “oui,” “ouais,” “ok,” or with “non,” the current utterance is tagged YN; if the utterance contains a French wh-word, it is tagged WH; if it contains “ou,” it is tagged DQ; otherwise it is tagged N/A. This automatic labeler assigned tags to 772 of 2,504 processed questions, approximately 31%, but systematic errors were observed, including WH questions followed by “non” being mis-tagged as YN and isolated “quoi ?” often being mis-tagged WH even when functioning phatically (Boritchev et al., 2022).

Human annotation involved 10 annotators, of whom 3 completed the full corpus and 7 annotated partial subsets. Full annotation took approximately 6 hours per annotator. Several operational decisions emerged during this process. For example, utterances ending in “ou pas” were tagged DQ because they contain “ou”; short forms such as “quoi ?,” “c’est bon ?,” “sérieusement ?,” and “encore ?” were generally treated as phatic unless context indicated otherwise; CS applied only when turn n1n-1 belonged to a different speaker, preventing split self-utterances from being mislabeled (Boritchev et al., 2022).

Across the three full annotators, the average question counts were: YN 1,441 (57.78%), WH 594 (23.82%), CS 8 (0.32%), DQ 97 (3.90%), PQ 304 (12.18%), and N/A 50 (1.99%). Inter-annotator agreement for the three full annotators was substantial: Cohen’s κ\kappa values were 0.651, 0.615, and 0.804 pairwise, and Fleiss’ κ\kappa was 0.693. In a first-half subset annotated by three partial annotators, Fleiss’ κ\kappa reached 0.813 (Boritchev et al., 2022).

The question layer makes DinG especially useful for pragmatic and interactional research. The corpus supports question detection, classification, turn-taking analysis, and comparisons between spontaneous human-human speech and more formal French question resources such as FQB. A plausible implication is that DinG occupies a distinctive niche between discourse corpora and spoken-language corpora: it is narrow in domain but dense in interactional phenomena (Boritchev et al., 2022).

5. Quality control, agreement, and ethical design

The corpus design is explicitly shaped by privacy and dissemination constraints. Names are removed, speakers are anonymized by colors, outside voices are labeled “Other,” and only transcriptions are released. The process was overseen by INRIA’s OCELER committee, complies with GDPR, and relies on retractable informed consent. The team is also exploring voice anonymization, including VoiceMask, for possible future release of audio without compromising privacy (Boritchev et al., 2022).

Transcription agreement reflects the inherent difficulty of noisy spontaneous dialogue. On a 5-minute excerpt of DinG2 with two independent annotators, κipf\kappa_{ipf} was 0.28 and raw agreement 0.28 when noise and pause durations were included. After removing noise tags and pauses, agreement improved to κipf=0.52\kappa_{ipf} = 0.52–0.53 with raw agreement 0.55. After super-annotation for consistency, κipf\kappa_{ipf} rose to 0.35 with raw agreement 0.35 (Boritchev et al., 2022).

The project therefore introduced a super-annotation stage that standardized typos, noise tags, onomatopoeias, numbers, and anonymization. This indicates that DinG is not merely a raw transcription dump but a curated resource with explicit normalization decisions. At the same time, the paper notes several limitations: the domain is constrained to a board-game context, recordings were made in noisy public settings, multi-threaded conversation complicates the alignment between questions and responses, and richer metadata such as token counts and vocabulary size are not provided (Boritchev et al., 2022).

These design choices address a common misconception about open spoken corpora. DinG is not open because privacy is absent; it is open because the interactional setting was selected to reduce personal disclosure and because anonymization and consent procedures were integrated into the corpus design from the outset (Boritchev et al., 2022).

6. Semantic extension through ding-01

The later ding-01 release extends DinG by adding AMR graphs to transcripts of spontaneous French dialogues recorded during Catan sessions. It preserves the original turn-taking segmentation from DinG and directly annotates French rather than projecting semantics from English resources or translations. The result is a semantic corpus tailored to spontaneous, multi-party French dialogue (Kang et al., 18 Aug 2025).

The annotated portion covers approximately 1,830 turns of speech, 1,667 non-empty utterances, 17,887 tokens, and 9 speakers. It includes 459 discourse markers and 36 backchannels. Train, development, and test splits for parser work are 1,375, 146, and 146 examples respectively, with non-annotable examples filtered out as needed. Each non-empty utterance has an AMR graph in PENMAN-style textual serialization (Kang et al., 18 Aug 2025).

ding-01 adheres closely to AMR 3.0 where possible but introduces minimal, removable extensions to represent phenomena that standard AMR underspecifies in spontaneous dialogue. These include the new roles :discourse-marker and :back-channel, cross-turn identifiers for inter-instance coreference and inter-instance verb ellipsis, discourse-sensitive root selection for clefts and left dislocation, and the use of :reparandum when a false start has interpretable semantic content (Kang et al., 18 Aug 2025).

The annotation workflow consisted of primary annotation by the first author over six months, with approximately 15% of examples validated by two co-authors and regular review meetings. Inter-annotator agreement on 160 doubly annotated examples, measured by Smatch, was 71.6. After disagreement resolution, the guidelines were refined to reduce conflicts, especially those involving :ARG0, :ARG1, :ARG2, and selection among near-synonymous PropBank concepts (Kang et al., 18 Aug 2025).

ding-01 also includes a parser baseline. Sequence-to-sequence AMR parsers based on mBART were trained in two regimes: a domain-specific setup trained solely on ding-01, and a pre-trained-plus-domain-specific setup pre-trained on French-translated AMR 3.0 and fine-tuned on ding-01. On the 146-example test set, Smatch scores were 68.1 for the domain-specific model and 73.5 for the pre-trained-plus-domain-specific model. Pre-training reduced ill-formed graphs and improved predicate selection, while both models captured :discourse-marker reasonably well, with around 30 of 43 instances found in the test set (Kang et al., 18 Aug 2025).

From the perspective of the original DinG corpus, ding-01 changes the resource’s role in the ecosystem. DinG is no longer only a transcription and dialogue-analysis corpus; it is also the substrate for a French semantic annotation framework aimed at spontaneous speech, dialogue pragmatics, and assistive semi-automatic AMR annotation (Kang et al., 18 Aug 2025).

7. Position in the French dialogue resource landscape

DinG fills a gap among French dialogue resources because many available corpora are either text-based, short, or sensitive, particularly in clinical or medical settings, which limits dissemination. Its proximity to STAC is especially significant: both resources are centered on Catan, but STAC contains English chat logs annotated in SDRT, whereas DinG contains oral French multi-party speech. This domain match with a modality mismatch creates a controlled basis for comparing written and oral discourse on the same task domain across languages and annotation frameworks (Boritchev et al., 2022).

The corpus is intended for dialog act and question classification, conversational analysis of turn-taking in multi-party settings, discourse-level modeling in SDRT or DRT-style work, and French ASR post-processing or transcription studies. The AMR extension adds semantic parsing, discourse modeling with graph structures, and graph querying through Grew (Boritchev et al., 2022, Kang et al., 18 Aug 2025).

Several limitations remain important. DinG is tied to a board-game context, which may limit lexical and topical diversity. Audio is not currently released. Token counts, average utterance length, and vocabulary size are not reported for the original corpus. In ding-01, details such as number of sessions, total duration of the annotated subset, or transcription policies for overlapping speech beyond the stated points are likewise not reported (Boritchev et al., 2022, Kang et al., 18 Aug 2025).

Even with those constraints, DinG has become a technically distinctive resource: an openly distributable corpus of approximately 11.7 hours of spontaneous French multi-party speech with time-aligned manual transcription, overlap representation, question-focused annotation, and an AMR extension designed specifically for spontaneous dialogue. That combination explains its relevance to both interactional linguistics and dialogue-oriented NLP (Boritchev et al., 2022)

Definition Search Book Streamline Icon: https://streamlinehq.com
References (2)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to DinG Corpus.