Modern Standard Arabic (MSA) Single-Utterance Speech Dataset
Commercially licensed Modern Standard Arabic speech data for ASR, speech foundation models, voice agents, and multilingual AI systems.
- Native Speakers
- Human-Verified
- Commercial License
- Production-Ready
SPX-AR-MSA-0001.wav
Female · 25–34
- Locale
- ar
- Audio
- 16 kHz · 16-bit · Mono
- Environment
- Quiet indoor
- Delivery
- WAV · JSONL
- Annotation
- Human transcript · Sentence-level alignment
One recording. One linked JSONL record.
Each WAV file maps to one JSONL record containing its transcript, speaker attributes, and technical metadata.
SPX-AR-MSA-0001.wav
{
"uniq_id": "SPX-AR-MSA-0001",
"duration": 29,
"language": "Arabic",
"text": "[Available in licensed delivery]",
"audio_path": "audio/SPX-AR-MSA-0001.wav",
"spkinfo": {
"language_code": "ar",
"spkid": "F0001",
"gender": "female",
"age_range": "25-34"
},
"sample_rate": 16000,
"bit_depth": 16,
"channels": 1,
"environment": "quiet_indoor",
"dataset_type": "single_utterance"
}WAV + Transcript + Metadata
Ready for delivery
Validated before delivery.
Every recording is reviewed across audio, transcription, and metadata before delivery.
- audio/*.wav
Audio
- 16 kHz · 16-bit PCM
- Signal integrity checked
- text
Transcript
- Human-reviewed
- Sentence-level alignment
- manifest.jsonl
Metadata
- Schema validated
- Speaker attributes included
- WAV + JSONLReady
Delivery
- WAV + JSONL
- Commercial licensing available
Records that fail validation are excluded from the delivery package.
Checked per record · Human reviewed
From licensed speech to production AI.
Structured audio, transcripts, and metadata support workflows across ASR training, pronunciation modeling, and benchmarking.
Automatic Speech Recognition
Train production Arabic ASR systems on aligned scripted read speech.
Speech Foundation Models
Support multilingual speech encoder and representation learning workflows.
Voice Agents
Improve Arabic voice interfaces, assistants, and conversational AI systems.
Model Evaluation
Benchmark recognition quality, pronunciation handling, and language coverage.
Deployed in production speech systems
Evaluate this dataset.
Review representative samples and documentation before licensing.
NDA available for qualified evaluations.