---
title: 'SwissGPC v1.0: Swiss German Podcasts Corpus'
url: https://www.emergentmind.com/topics/swissgpc-v1-0
type: topic
---

# SwissGPC v1.0: Swiss German Podcasts Corpus

Searching arXiv for the named resource and paper metadata.
SwissGPC v1.0 is the Swiss German Podcasts Corpus, described as the first mid- to large-scale corpus of spontaneous Swiss German speech and intended to support research in ASR, TTS, dialect identification, speech translation between Swiss German and Standard German, speaker-related modeling, and other weakly supervised spoken-language tasks [2509.19866]. Its distinguishing features are scale, spontaneous conversational speech, and a fully automated weak-annotation pipeline applied to publicly available podcasts and talk-show material from Schweizer Radio und Fernsehen and YouTube. Unlike earlier Swiss German corpora that are smaller and largely based on scripted, elicited, or otherwise controlled recordings, SwissGPC emphasizes natural conversations with hesitations, interruptions, informal turn-taking, and topic shifts, thereby targeting conditions closer to real deployment.

## 1. Corpus identity and research setting

SwissGPC v1.0 was motivated explicitly by the need to train a zero-shot voice adaptation text-to-speech system for Swiss German dialects, but the resource is positioned more broadly as usable for automatic speech recognition, dialect identification, speech translation, speaker-related modeling, and weakly supervised speech tasks [2509.19866]. Its central design choice is to privilege spontaneous, long-form audio over carefully controlled recording conditions. This places it in a different methodological category from corpora such as Swiss Parliament Corpus, SwissDial, SDS-200, and STT4SG-350, which the paper characterizes as substantially smaller and mostly based on controlled speech.

The scale contrast is one of the paper’s main claims. SwissGPC draws on 5,404 hours of raw audio and retains 4,979.09 hours of speech after processing. By comparison, the earlier Swiss German corpora cited in the paper range from 28 to 343 hours. This suggests a shift from corpus construction oriented around annotation cleanliness and balanced elicitation to corpus construction oriented around breadth, realism, and compatibility with data-hungry models.

A common misconception would be to treat SwissGPC v1.0 as a conventional packaged speech dataset with redistributed media and gold labels. The paper states the opposite. SwissGPC is released as a list of source links and code for downloading audio and running the automated annotation pipeline, while neither audio files nor annotated data from SRF and YouTube are redistributed because of copyright restrictions. In practical terms, it is a reproducible corpus recipe rather than a conventional download bundle.

## 2. Source collection and corpus composition

The corpus is assembled from long-form public audio hosted on Schweizer Radio und Fernsehen and YouTube, specifically podcasts and talk-show-like programs rather than isolated utterances [2509.19866]. On the SRF side, 25 podcast sources contribute 4,715.87 hours of raw audio. The paper lists major contributors including *Tagesgespräch* with 1,661.33 hours, *Digital Podcast* with 428.05 hours, *Samstagsrundschau* with 404.14 hours, *Wissenschaftsmagazin* with 393.61 hours, *Geek-Sofa* with 317.28 hours, and *Debriefing 404* with 245.14 hours. On the YouTube side, 12 sources contribute 688.93 hours, including *Feel Good Podcast* with 319.60 hours, *Berner Jugendtreff* with 127.80 hours, *Finanz Fabio* with 58.44 hours, and *Fadegrad* with 49.95 hours.

The raw collection comprises 15,171 individual episodes. The average episode length is 1,277.28 seconds, or 21.28 minutes. The duration distribution is highly uneven: one peak lies around 100–200 seconds, attributed largely to *100 Sekunden Wissen*, and another around 1,600–1,800 seconds, corresponding to the common 20–30 minute podcast format. There are 32 outlier episodes longer than \(7200s\); the longest lasts 13,846 seconds, or 3 h 50 min, and the shortest lasts 19 seconds. The median number of episodes per podcast is 104, while the mean is 410, indicating strong skew. The largest single source, *Tagesgespräch*, contributes nearly 31% of the total raw dataset.

After processing, 4,979.09 hours of speech are retained, representing a reduction of 7.84% from the raw audio. This retained material is segmented into 1.767 million samples, reported elsewhere in the paper as 1.76M and 1767K. The token count is reported inconsistently: the main text gives 55.85M tokens, while the dialect statistics table sums to 55.81M. The paper does not explain the discrepancy, so both figures remain paper-reported values rather than reconciled corpus totals.

No train/dev/test split is provided for SwissGPC v1.0. The paper also does not report a total speaker count, speaker demographics, or a speaker-balanced partitioning scheme. These omissions matter for comparative benchmarking and for experiments requiring explicit speaker controls.

## 3. Automated construction and weak annotation pipeline

SwissGPC v1.0 is built with a fully automated, weakly supervised pipeline [2509.19866]. The workflow begins by gathering public podcast and talk-show content from SRF and YouTube. For SRF, collection relies on the official SRG SSR API; for YouTube, the paper names the `pytubefix` fork of PyTube.

The audio is then diarized with `pyannote.audio`. Because this step marks speech regions and speakers, silence and music-only regions are implicitly removed. The diarized speech is segmented into clips constrained to be between 2 and 15 seconds long, and the segments are intended to contain only a single speaker according to the diarization output. The paper presents this duration range as a compromise that preserves variability while also meeting downstream model requirements for transcription and training.

Each segment is transcribed into Standard German with Whisper v3. The choice of Standard German rather than Swiss German is deliberate: Standard German has standardized orthography, whereas Swiss German does not. A wav2vec2-based phoneme transcriber is then used to derive phoneme sequences. These phoneme sequences are classified by a Naïve Bayes \(n\)-gram model trained on phonemicized STT4SG-350 data plus a Standard German CommonVoice subset in order to assign a dialect or language-region label. The paper also describes optional enrichment steps: generating Swiss German text with a pipeline from prior work and computing mel-spectrograms with `librosa`.

The resulting annotations are weak rather than gold. They include segment boundaries and timestamps implied by diarization and cutting, single-speaker segmentation, Standard German transcripts from Whisper v3, dialect-region labels predicted by the Naïve Bayes classifier, generated Swiss German text, and mel-spectrogram features. Speaker separation is available implicitly at the segment level because of diarization, but the paper does not state that persistent speaker identities across episodes are released.

The paper does not provide explicit equations, loss functions, or confidence thresholds for the diarization model, Whisper decoding, phoneme recognizer, or Naïve Bayes classifier. The explicit operational thresholds that are given are the 2–15 second segment bounds and the descriptive outlier threshold of \(>7200s\) for episode-length plots. It also observes that very short samples \((< 7\) tokens) and very large samples \((\geq 65\) tokens) were often erroneous or incoherent Whisper translations, but presents this as a descriptive observation rather than as a formal released filtering rule.

## 4. Linguistic coverage and dialect labeling

SwissGPC covers the seven major Swiss German dialect regions used in STT4SG-350, together with Standard German as an eighth class [2509.19866]. The seven regions are Basel, Bern, Central CH, Eastern CH, Grisons, Valais, and Zurich. In the dialect-identification setup, the classifier therefore predicts one of eight labels: Basel, Bern, Central CH, Eastern CH, Grisons, Valais, Zurich, or Standard German.

The inclusion of Standard German is a structural consequence of the source material. Many programs, especially more formal SRF content on science, philosophy, news, or public affairs, are delivered in Standard German rather than Swiss German. The paper also notes an unresolved ambiguity: it cannot distinguish non-Swiss Standard German from Swiss speakers using Swiss Standard German pronunciation, so all such material is grouped into a single “German” label. This creates a linguistically heterogeneous class, though the authors hypothesize that such material may still be useful for training speech models.

Dialect labels are effectively assigned at the segment level rather than only at the podcast or speaker level. The pipeline diarizes audio into single-speaker segments, converts each segment to a phoneme sequence, and classifies that sequence into one of the eight region or language labels. The paper’s validation on *Zivadiliring* also describes segment-level predictions that align with different hosts and with a Basel guest. At the same time, the paper does not claim speaker-level gold dialect labels or manual utterance-level labels.

The distribution is highly imbalanced. The paper explicitly describes the corpus as “highly unbalanced on a dialectal basis.” The following table reproduces the reported hours and token counts.

| Region | Hours | Tokens |
|---|---:|---:|
| Basel | 460.81 h | 5.35M |
| Bern | 771.38 h | 8.98M |
| German | 1685.72 h | 17.23M |
| Grisons | 151.33 h | 1.74M |
| Central CH | 341.22 h | 3.95M |
| Eastern CH | 350.60 h | 4.00M |
| Valais | 39.46 h | 0.43M |
| Zurich | 1178.58 h | 14.13M |
| Total | 4979.09 h | 55.81M |

By sample count, the same table reports Basel 179K, Bern 293K, German 538K, Grisons 57K, Central CH 121K, Eastern CH 121K, Valais 15K, and Zurich 440K, totaling 1767K samples. The two largest classes, Standard German and Zurich, together account for 57.53% of the corpus duration, while Valais constitutes only 0.79%. The paper further notes that Standard German segments tend to contain more tokens, likely because formal informational programs are denser and more monologic.

## 5. Acoustic and structural characteristics

SwissGPC is designed around spontaneous, natural, uncontrolled speech in generally high-quality recordings [2509.19866]. The choice of podcasts reflects a practical assumption that such material offers good recording conditions together with speaker and topic diversity. The paper does not, however, provide technical recording specifications such as sample rate, bit rate, microphone types, or channel format.

The corpus is meant to capture properties often absent from read-speech datasets: hesitations, interjections, interruptions, overlapping speech, and informal turn structure. Overlap handling is indirect rather than explicitly parameterized. Because segments are derived from diarization output and are supposed to contain one speaker, overlap is plausibly reduced or split by the diarizer, but the paper does not quantify overlap rates or present a dedicated overlap-resolution heuristic.

Several structural properties follow from segmentation and automatic transcription. Segment token distributions are bimodal. One peak lies around 7–14 tokens, attributed to short speaker turns and rapid turn-taking in spontaneous podcast speech. A second peak lies around 40–53 tokens, attributed to longer and denser monologues such as storytelling or reading. No mean or median segment duration is reported directly. The paper also observes that training a TTS model on these segments led to better generation on longer segments than on shorter ones, suggesting that segment granularity has downstream consequences. This suggests, though does not prove, that weak segmentation decisions interact materially with synthesis quality.

The transcription regime is mixed. Standard German transcripts are automatic outputs from Whisper v3 and serve as the main textual supervision signal because Standard German orthography is standardized. Swiss German transcripts are also generated in an enrichment step, but the paper characterizes them as much noisier because Swiss German lacks standardized orthography. No custom orthographic normalization scheme is introduced. The paper’s examples show that generated Swiss German forms can diverge substantially from manual references even when content remains semantically similar.

## 6. Validation, use cases, and constraints

The paper validates several pipeline components on the *Zivadiliring* podcast, chosen because it is around 50 hours, exclusively Swiss German, has limited guest variation, and includes hosts from known dialect regions: one from Eastern Switzerland and two from Zurich [2509.19866]. Diarization is manually evaluated on a 42 minute 38 second episode using ELAN, yielding a diarization error rate of 14.1%. The paper compares this with 12.2% on AISHELL-4 for the same pipeline.

Standard German transcription quality is evaluated on 100 randomly selected audio samples and reaches a word error rate of \(0.30(\pm 0.264)\). The paper attributes the remaining errors to systematic differences between Swiss German and Standard German, including tense usage, auxiliary verbs, grammar, Helvetisms, loanwords, and information loss when dialect speech is translated into Standard German. On the same 100 samples, Swiss German transcription yields a much higher WER of \(0.639(\pm 0.253)\), which the authors mainly attribute to the lack of standardized orthography.

The dialect classifier is retrained with an added Standard German class using 30 hours from CommonVoice matched for age and gender distribution to the STT4SG-350 training material, and achieves macro F1 \(= 0.88\) across eight regions. On *Zivadiliring*, about two-thirds of the segments are labeled Zurich and one-third Eastern CH, which aligns with the hosts’ origins; a Basel guest is also correctly identified in a specific episode. These evaluations support usability of the weak-annotation pipeline while also confirming that the annotations are noisy rather than gold-standard.

The paper does not report full downstream benchmarks trained directly on SwissGPC, such as ASR or TTS baselines with published train/dev/test protocols and scores. It mentions one use case, voice adaptation for Swiss German dialects with XTTSv2, and states that a weakly supervised approach on this type of data yielded very good results, but it does not provide detailed setup or metrics in the paper. Likewise, although the corpus is framed as suitable for ASR and speech translation, there is no standard benchmark section for SwissGPC itself.

Access and reproducibility are constrained by licensing. The released artifact consists of source links and code for downloading podcasts and running the pipeline; audio and annotated data are not distributed because legal rights are not held for SRF and YouTube content. The authors describe the corpus as a “snapshot in time,” and warn that exact reproduction may become difficult because episodes may be added, renamed, moved, geo-restricted, or removed. Versioning is limited to “v1.0,” with no semantic version policy or frozen archival mirror defined. The paper also notes several other limitations: noisy automatic annotations, strong dialect imbalance, incomplete source coverage, ambiguity inside the Standard German class, absence of speaker metadata, and dependence on third-party platform availability. Future work is suggested in the form of broader source crawling, especially for under-represented dialect regions, inclusion of additional SRF podcasts and television shows, and possible improvements to segmentation and transcription.

Source: https://www.emergentmind.com/topics/swissgpc-v1-0