G-SODA is our proprietary collection of large-scale multilingual speech corpora, developed and curated since the 1980s.
Trusted by AI researchers and developers worldwide, these datasets support the training and evaluation of speech recognition systems, conversational AI, spoken language technologies, and next-generation foundation models.
Explore selected samples from our professionally recorded, studio-quality corpora below.
Sample Speech Data
Note 1: The audio samples below have been downsampled to 16 kHz, 16-bit for web playback. The original product is delivered in 48 kHz, 32-bit (or 24-bit / 16-bit, depending on the product specifications).
Note 2: Each sample below is a short excerpt from a single audio file. Most of our conversational speech recordings are approximately 50 minutes in length per file.
Sample 8. Japanese Native Speech – Multi-Party Conversation (CCSJJ)
Sample 9. Japanese Native Speech – Telephone Conversation (CPSJJ)
Sample 12. Japanese Native Speech – Medical Conversations (CHSJJ)
About These Speech Samples
Sample 1. French (Martinique × Réunion) – Spontaneous Conversation
Language: French
Speakers: Native speakers from Martinique (French Caribbean) and Réunion Island (French Indian Ocean)
Speech Style: Spontaneous conversation
A naturally occurring French conversation between speakers from two geographically distinct French-speaking regions. This sample illustrates regional variation within the French-speaking world while preserving authentic conversational speech.
Sample 2. Chinese (Hong Kong) – Spontaneous Conversation
Language: Chinese
Speakers: Native speakers from Hong Kong
Speech Style: Spontaneous conversation
A spontaneous conversation demonstrating naturally occurring speech patterns and regional characteristics commonly found among speakers from Hong Kong.
Sample 3. Hong Kong English – Spontaneous Conversation
Language: English
Speakers: Native English speakers from Hong Kong
Speech Style: Spontaneous conversation
A natural conversation between native English speakers from Hong Kong. This sample demonstrates a regional variety of English shaped by a multilingual and international linguistic environment.
Sample 4. Cantonese – Spontaneous Conversation
Language: Cantonese
Speakers: Native Cantonese speakers
Speech Style: Spontaneous conversation
A spontaneous conversation between native Cantonese speakers. Cantonese is a major Sinitic language spoken in Hong Kong, Macau, and southern China, and is linguistically distinct from Mandarin Chinese.
Sample 5. Japanese (Non-Native Speakers from Hong Kong) – Conversation
Language: Japanese
Speakers: Native Hong Kong speakers using Japanese as a foreign language
Speech Style: Spontaneous conversation
A natural Japanese conversation between non-native speakers from Hong Kong. The recording captures authentic second-language speech characteristics, including pronunciation, fluency, and learner-specific communication patterns.
Sample 6. Japanese Native Speech – Monologue (CMSJJ)
Language: Japanese
Speakers: Native Japanese speakers
Speech Style: Monologue
A continuous monologue produced by a native Japanese speaker without interaction from other participants. This sample is suitable for speech recognition, speaker modeling, and spoken language research.
Sample 7. Japanese Native Speech – Dialogue (CDSJJ)
Language: Japanese
Speakers: Native Japanese speakers
Speech Style: Two-person dialogue
A natural dialogue between two native Japanese speakers. The recording contains authentic turn-taking behavior and interactive conversational speech.
Sample 8. Japanese Native Speech – Multi-Party Conversation (CCSJJ)
Language: Japanese
Speakers: Native Japanese speakers
Speech Style: Three-person conversation
A spontaneous multi-party conversation featuring three native Japanese speakers. The recording captures overlapping speech, turn transitions, interruptions, and natural conversational dynamics.
Sample 9. Japanese Native Speech – Telephone Conversation (CPSJJ)
Language: Japanese
Speakers: Native Japanese speakers
Speech Style: Telephone conversation
A naturally occurring telephone conversation between native Japanese speakers. The sample reflects real-world remote communication conditions and telephone speech characteristics.
Sample 10. Japanese Native Speech – University Lecture (CLSJJ)
Language: Japanese
Speakers: Native Japanese speakers
Speech Style: Academic lecture
A university lecture delivered by a native Japanese speaker. This sample contains long-form educational speech produced in an academic environment.
Sample 11. Japanese Native Speech – Children's Speech (CKSJJ)
Language: Japanese
Speakers: Native Japanese-speaking children
Speech Style: Children's speech
A speech sample recorded from native Japanese-speaking children. The corpus captures age-dependent speech characteristics, pronunciation patterns, and spontaneous child language behavior.
Sample 12. Japanese Native Speech – Medical Conversations (CHSJJ)
Language: Japanese
Speakers: Native Japanese speakers
Speech Style: Medical and healthcare conversations
Authentic conversations recorded in medical and healthcare settings. The sample includes interactions between patients, healthcare professionals, and related participants.
Sample 13. Japanese Native Speech – Read Speech (CTSJJ)
Language: Japanese
Speakers: Native Japanese speakers
Speech Style: Read speech
A speech sample consisting of native Japanese speakers reading novels, stories, and other written materials. The recording captures carefully articulated read speech under controlled conditions.
Timehill Inc

