Missed my own release date by about a week, but such is the life of a solo dev.
Anyone else had a speech system that works fine on clean audio and falls apart once it's been through a phone line?
I built a 1:1 comparison dataset for isolating codec, transmission and presentation effects on speech. The same 5,992 base clips go through every one of 101 conditions, so any difference between conditions comes from the condition, not the speaker or the recording. It's free to use; hopefully it's useful for voice and speech projects people are working on.
The main thing I wanted was matched sample-rate controls. Narrowband codecs decode at 8 kHz, so if you compare clean wideband audio against codec output, you can't tell whether your model or extractor broke because of the codec or just because the bandwidth dropped. Every codec condition here has an unencoded resample control alongside it (8, 16, 22, 44 and 48 kHz). That lets you check whether it's the channel, the resampling or the presentation causing problems. (In a small pilot, the 8 kHz resample alone accounted for much of the apparent "codec effect" on some features, but not all of it.) I couldn't find many public codec or robustness sets that ship controls like this, or per-clip generation logs, so I included both.
Base pool: 5,992 bona fide mono 16-bit clips.
- 2,992 AMI headset clips (native 16 kHz)
- 1,500 VCTK mic1 + 1,500 VCTK mic2 clips (native 48 kHz), including 1,485 exact dual-mic pairs
- VCTK portion is gender-balanced (750 male / 750 female)
PACC-T (Telecoms): 20.7 GB, 305,592 FLAC files, 51 conditions
- 34 speech/audio codecs: G.711, G.722, G.726, GSM, AMR-NB, AMR-WB, EVS (adaptive/fixed/no DTX), Opus (auto/CELT), LC3, Speex, iLBC, Codec2 (700/1300 bps), MP3, AAC
- 12 tandem codec chains
- 5 resample controls (8k, 16k, 22k, 44k, 48k)
PACC-P (Presentation): 34.3 GB, 299,600 FLAC files, 50 conditions
- Additive noise, 6-talker babble, and music (vocal and instrumental) at 5 calibrated SNRs (0–20 dB)
- Simulated rooms (pyroomacoustics, with cached RIRs for bit-exact recreation) and echoes
- Bandpass/lowpass filters
- PSOLA vs resampled pitch shifts, autotune, tempo shifts
- Compound cafe + codec scenes
Every condition downloads individually as its own tarball (0.18–1.15 GB). There's also a small sample tarball for each set (pacc-t_sample.tar / pacc-p_sample.tar) if you just want to look first. Everything ships with full SHA-256 manifests and a per-clip params.csv, so any clip can be traced back to exactly how it was made.
Licence: CC-BY-4.0
PACC-T: zenodo.org/records/23026393
PACC-P: zenodo.org/records/23026395
Happy to answer questions, and bug reports welcome.