Back to catalog
Commercial Speech DatasetAvailable for licensing

Danish Single-Utterance

Danish single-utterance speech from native speakers, recorded from scripted prompts for ASR training, speech model adaptation, pronunciation modeling, and evaluation.

Request a sample
  • Native Speakers
  • Human-Verified
  • Commercial License
  • Production-Ready
Sample Audioda-DK

SPX-DA-DK-0001.wav

Female · 25–34

0:00--:--
Audio Samples4 representative recordings
Locale
da-DK
Audio
16 kHz · 16-bit · Mono
Environment
Quiet indoor
Delivery
WAV · JSONL
Annotation
Human transcript · Sentence-level alignment
Record Anatomy

One recording. One linked JSONL record.

Each WAV file maps to one JSONL record containing its transcript, speaker attributes, and technical metadata.

Audio Record

SPX-DA-DK-0001.wav

0:0616 kHz · 16-bit · Mono
linked by uniq_id
Linked JSONL Record
manifest.jsonlSPX-DA-DK-0001
{
  "uniq_id": "SPX-DA-DK-0001",
  "duration": 6,
  "language": "Danish",
  "text": "[Available in licensed delivery]",
  "audio_path": "audio/SPX-DA-DK-0001.wav",
  "spkinfo": {
    "language_code": "da-DK",
    "spkid": "SPK-707",
    "gender": "female",
    "age_range": "25-34"
  },
  "sample_rate": 16000,
  "bit_depth": 16,
  "channels": 1,
  "environment": "quiet_indoor",
  "dataset_type": "single_utterance"
}

WAV + Transcript + Metadata

Ready for delivery

Quality Assurance

Validated before delivery.

Every recording is reviewed across audio, transcription, and metadata before delivery.

  1. audio/*.wav

    Audio

    • 16 kHz · 16-bit PCM
    • Signal integrity checked
  2. text

    Transcript

    • Human-reviewed
    • Sentence-level alignment
  3. manifest.jsonl

    Metadata

    • Schema validated
    • Speaker attributes included
  4. WAV + JSONLReady

    Delivery

    • WAV + JSONL
    • Commercial licensing available

Records that fail validation are excluded from the delivery package.

Checked per record · Human reviewed

Use Cases

From licensed speech to production AI.

Structured audio, transcripts, and metadata support workflows across ASR training, pronunciation modeling, and benchmarking.

This dataset · WAV + JSONL
  • ASR Training

    Train Danish speech recognition systems with clean, human-transcribed read speech.

  • Pronunciation Modeling

    Improve Danish phoneme, pronunciation, and acoustic modeling.

  • Benchmarking

    Evaluate recognition quality on controlled, sentence-level Danish speech.

Deployed in production speech systems

Next Step

Evaluate this dataset.

Review representative samples and documentation before licensing.

Request a sample

NDA available for qualified evaluations.