CIMA Corpus: Cross-Lingual Meeting Data
- CIMA Corpus is a comprehensive dataset of cross-lingual meeting dialogues featuring simultaneous speech translation, multitrack audio, and incremental minuting.
- It includes detailed artifacts such as ASR outputs, human-corrected transcripts, English minutes, and explicit misunderstanding annotations for versatile evaluation.
- The corpus supports research on real-time translation, cross-lingual summarization, and automatic misunderstanding detection through its realistic, interactive meeting recordings.
The CIMA Corpus, described in “Corpus of Cross-lingual Dialogues with Minutes and Detection of Misunderstandings,” is a corpus of cross-lingual dialogues between individuals without a common language who were facilitated by automatic simultaneous speech translation, with additional artifacts for minuting and misunderstanding analysis (Čechovič et al., 23 Dec 2025). In the paper and release, the corpus itself is named InCroMin—Incremental Cross-Lingual Minutes—reflecting its focus on meetings where translation and minutes are produced incrementally and in real time. Its defining contribution is the combination of realistic, interactive, machine-mediated multilingual meetings with multi-track audio, ASR outputs, corrected transcripts and translations, English minutes, and explicit misunderstanding annotations, thereby supporting evaluation of simultaneous speech translation, cross-lingual summarization, and misunderstanding detection in a single resource (Čechovič et al., 23 Dec 2025).
1. Nomenclature and research setting
Within the 2025 corpus paper, “CIMA Corpus” denotes the resource presented under the release name InCroMin (Čechovič et al., 23 Dec 2025). The nomenclature matters because the release foregrounds incremental minuting, whereas the article title foregrounds cross-lingual dialogues, minutes, and misunderstandings. In practice, both labels refer to the same resource in this context.
The corpus was motivated by the need for a versatile and realistic evaluation corpus for meetings of individuals who do not share any common language. This motivation differentiates it from read speech or monologue sentence collections that dominate prior resources. The paper explicitly contrasts the corpus with datasets such as LibriSpeech and MuST-C, emphasizing spontaneity, interactivity, the lack of sentence boundaries, turn-taking issues, delayed translations, and technical artifacts as core characteristics of the setting (Čechovič et al., 23 Dec 2025). This suggests that the corpus is designed less as a controlled benchmark for isolated utterance translation than as an integrated testbed for end-to-end machine-mediated conversation.
Meetings were held in an online videoconference setting. Participants deliberately used different primary languages and relied on machine-mediated communication rather than reverting to a shared language such as English. The resulting dialogues are therefore not merely multilingual recordings; they are interactions structured around simultaneous translation latency, interface constraints, and the possibility of communication breakdowns (Čechovič et al., 23 Dec 2025).
2. Collection design and corpus composition
The first public release contains 10 meetings totaling 296 minutes, approximately 5 hours of audio including silences, with 24 participants across 12 original languages (Čechovič et al., 23 Dec 2025). Meetings involved 2–5 participants: one meeting with 5 participants, one with 3 participants, and the rest with 2. Participants were volunteers drawn from the authors’ networks and were matched so that at least two groups in each meeting used different languages that were not mutually understood. They were asked not to revert to any common language they might share.
Two meetings simulate interviews between a refugee and an integration center officer, conducted by two officers to resemble real interactions, because authentic refugee data collection with consent is typically infeasible (Čechovič et al., 23 Dec 2025). Meetings were intended to last 30 minutes, followed by a feedback questionnaire of about 15 minutes.
| Component | Released value | Notes |
|---|---|---|
| Meetings | 10 | First public release |
| Duration | 296 minutes | About 5 hours including silences |
| Participants | 24 | 2–5 per meeting |
| Original languages | 12 | No participant selected English as a primary language |
| Tokens | 26,854 | From corrected transcripts where available or ASR otherwise |
The 12 languages in the released portion are Czech, Russian, Chinese, Brazilian Portuguese, Slovak, French, Italian, Ukrainian, Vietnamese, Marathi, Portuguese, and Spanish (Čechovič et al., 23 Dec 2025). English functions as the hub target language: Whisper Large-v3 translated from the original languages into English, and NLLB-200 distilled was used to translate English back into participants’ preferred languages during meetings. Language distribution is unbalanced, with Czech representing 45% of the audio due to the availability of volunteers.
A planned future release will add another 10 meetings, 20 participants, and 4 additional languages, for a total of 596 minutes (Čechovič et al., 23 Dec 2025). The public corpus handle is http://hdl.handle.net/11234/1-5956.
3. Collection platform, processing pipeline, and released artifacts
Meetings were held online via Fairmeeting, with the Minuteman tool facilitating simultaneous speech translation, minuting, and data collection (Čechovič et al., 23 Dec 2025). Minuteman recorded separate audio tracks for each participant, providing implicit speaker identification and removing the need for diarization. This is a consequential design decision because it yields multi-speaker meeting data without a downstream diarization dependency.
The reproducibility description enumerates the post-processing steps: retrieval of per-participant audio and displayed translation logs from Minuteman; synchronization and trimming to meeting bounds; ASR and English speech translation; human corrections; deidentification; minutes creation; and misunderstanding annotation (Čechovič et al., 23 Dec 2025). The paper does not specify hardware or audio sampling rates and formats.
The released corpus contains several aligned artifact layers. Multi-track audio is synchronized and trimmed to meeting content. ASR outputs and English speech translations are produced automatically, with human corrections added where possible. In the first release, 15 of 24 transcripts, or 63%, and 13 English translations, or 54%, are corrected. The resource also includes logs of displayed translations aligned to meeting time and trimmed to match the audio (Čechovič et al., 23 Dec 2025).
English minutes are included as written summaries created by an annotator. The paper defines “minutes” as brief written records summarizing agenda, conclusions, and action items. During meetings, Minuteman could translate minutes to participants’ languages, but the released annotation is in English. The paper positions these minutes as gold summaries for cross-lingual summarization research but does not report summarization baselines, dataset splits, or metrics such as ROUGE or BERTScore (Čechovič et al., 23 Dec 2025).
Deidentification is performed by removing names, locations, and organizations for private persons from text and replacing the corresponding audio segments with silence, with reported time intervals of removals. An annotator proficient in the spoken language performed this step. Silero VAD was used to mark voiced segments and silence, and SacreMoses was used for tokenization, including morphological segmentation for non-whitespace languages such as Chinese (Čechovič et al., 23 Dec 2025).
4. Minutes, summarization, and misunderstanding annotation
The minutes make the corpus directly relevant to cross-lingual meeting summarization. The specific emphasis is not generic abstractive summarization but automatic minuting in a setting where translations and minutes grow incrementally and in real time (Čechovič et al., 23 Dec 2025). A plausible implication is that the corpus supports research on temporally grounded summarization under translation latency and speaker alternation, rather than only post hoc summarization from fully corrected transcripts.
The misunderstanding layer is the most distinctive annotation contribution. The paper proposes a novel annotation protocol for misunderstandings in cross-lingual meetings, using revised English translations as the working text because annotators are not proficient in all original languages (Čechovič et al., 23 Dec 2025). A misunderstanding is operationalized as a segment, termed a “markable,” where some miscommunication occurred, typically spanning several sentences.
Each misunderstanding is annotated along three dimensions. First, annotators select markable spans. Second, awareness is labeled as “acknowledged” or “unrecognized,” depending on whether the speakers promptly display awareness. Third, the misunderstanding is assigned a reason, attributed to speaker A or B: translation error, delay, technical problem, or genuine misunderstanding. The last category refers to non-technological causes such as background knowledge gaps or ambiguous references, that is, errors likely to arise even in monolingual dialogues (Čechovič et al., 23 Dec 2025).
Across 14 two-participant meetings from both the first and planned releases, the annotation process produced 25 annotations and 222 misunderstandings in total (Čechovič et al., 23 Dec 2025). The distribution of reasons is 36.5% translation errors, 29.7% genuine misunderstandings, 14.4% delays, and 19.4% technical problems. Of all misunderstandings, 57% were acknowledged and 43% were unrecognized. The paper also contrasts meetings with Czech versus without Czech, hypothesizing shared background knowledge in Czech meetings leads to fewer translation and genuine misunderstandings and higher awareness.
5. Agreement methodology and automated detection
Eleven meetings were double-annotated. Rather than using item-level matching, inter-annotator agreement is assessed via Pearson correlation of counts per category per meeting (Čechovič et al., 23 Dec 2025). The reported coefficients are 0.94 for bad translation, 0.72 for technical problem, 0.83 for acknowledged misunderstandings, 0.90 for unrecognized misunderstandings, 0.55 for genuine misunderstanding, and 0.49 for delay. Annotators disagreed notably on “Gemini false positives,” with a correlation of –0.13, traced to a handful of meetings where one annotator rejected several Gemini suggestions while the other rejected none.
Automatic misunderstanding detection was tested with Gemini 1.5 Pro in a private Google Cloud environment using a concise prompt instructing the model to identify misunderstandings and provide timestamps and reasoning (Čechovič et al., 23 Dec 2025). Annotators then marked the model’s findings as true positive or false positive. The reported results are recall of 77% and precision of 47%.
Using the paper’s formulas,
the corresponding F1 score is approximately 0.584, or about 58.4% (Čechovič et al., 23 Dec 2025). The paper explicitly cautions that the dataset is small and the false-positive characterization is limited, so generalization should be avoided. This caution is central: the result is presented as an initial evaluation with a LLM, not as a definitive benchmark for robust misunderstanding detection.
6. Limitations, ethical considerations, and disambiguation
Several limitations are structural rather than incidental. The language distribution is opportunistic, with many Czech speakers; genuine cross-lingual need was simulated, because participants often could speak English but were asked not to; and two meetings are domain simulations in refugee support (Čechovič et al., 23 Dec 2025). The recording environment is online conferencing, so the corpus inherits latency, segmentation, and UI effects that are important for realism but also complicate controlled evaluation.
The paper identifies multiple observed failure modes: wrong target language selection, errors due to voice activity detection and segmentation, incorrect translation of dialogue-management meta-language such as turn-taking questions, mistranslation of sentence modality, latency-induced turn-taking problems, and unclear output presentation regarding whether more translation is yet to come (Čechovič et al., 23 Dec 2025). Participants reported a noticeable slowdown, often around compared to normal calls, and variable fatigue. Requested improvements included better synchronization, clearer speaker separation, and UI enhancements such as push-to-talk.
Ethically, voice is personally identifiable; informed consent was obtained; and deidentification removes private names and entities while silencing corresponding audio where necessary (Čechovič et al., 23 Dec 2025). The paper also identifies the phenomenon under study itself as an ethical risk: miscommunication when automatic systems mediate human interaction.
The term “CIMA” is also potentially ambiguous in the broader literature. Unrelated uses include the CPS mitigation system “Countering Illegal Memory Accesses” (Chekole et al., 2018), an educational dialogue setting in which a later paper refers to the open-source CIMA corpus as “Conversational Instruction with Multi-responses and Actions” (He et al., 11 Sep 2025), and an astrophysics usage for a C II morpho-kinematics corpus (Flynn, 25 May 2026). Conversely, “CEMA” is a distinct medieval charter meta-corpus rather than a CIMA corpus (Perreaux, 2021), and the Contemporary Amharic Corpus is designated CACO, not CIMA (Gezmu et al., 2021). In the specific sense established by the 2025 corpus paper, however, CIMA refers to the cross-lingual meeting corpus released as InCroMin (Čechovič et al., 23 Dec 2025).