Arabic (Saudi Arabia) Spontaneous Dialogue
This dataset contains spontaneous two-party conversations recorded by native Saudi Arabic speakers across Central, Western, and Eastern regional dialects. Human-verified transcripts with utterance-level timestamps support production ASR, automotive voice systems, voice agents, and multilingual speech model development.
Dataset Overview
Commercial License
Cleared for production AI development and model training.
Native Speakers
Native Saudi speakers covering Central · Western · Eastern Saudi Arabia.
Human Verified
Human transcription with dialect notes, reviewed before delivery.
Production Ready
16 kHz / 16-bit PCM WAV with structured, validated metadata.
Hear the data before you license it.
A representative audio excerpt from this dataset. Full transcripts and speaker metadata are delivered with licensed data.
Audio Sample 1
25–34 · Saudi Arabic
Built to a documented standard.
Every delivery matches the recording, audio, and annotation specification below.
Recording
Two-party conversational recordings in quiet indoor environments across several regional dialects.
Audio
Annotation
- Human Transcript
- Speaker ID
- Timestamps
- Dialect Tags
- JSONL / TXT
Utterance-level timestamps
Balanced, documented speaker coverage.
Native Saudi speakers
Balanced
18–60
Central · Western · Eastern Saudi Arabia
Verified at every stage of delivery.
What teams build with this dataset.
Automatic Speech Recognition
Train and benchmark ASR models against human-verified reference transcripts.
Voice Agents
Build voice assistants that hold up against real conversational speech.
Multilingual AI Products
Extend language coverage with consistent annotation across locales.
License this dataset for production AI.
Commercial licensing for production AI development. Terms are shared on request; evaluate representative samples before purchase.