Audio input

16 kHz mono 16-bit WAV is read directly. Any other format is decoded by shelling out to ffmpeg, which must be installed and on PATH. The decoded audio is staged in a temporary file (removed afterwards), so decoding needs write access to the system temp directory.

Verified to decode, and covered by tests:

kind tested
audio containers/codecs WAV, MP3, Ogg (Vorbis), Opus, M4A (AAC), ADTS AAC, FLAC
sample rates 8 kHz, 16 kHz, 44.1 kHz, 48 kHz, 96 kHz
bit depths 16-bit, 24-bit, 32-bit float
channel counts mono, stereo, 8-channel
video containers MP4 (H.264+AAC), MKV (H.264+AC3), WebM (VP9+Opus)

Video input is read for its audio track; a file with no audio track is reported as one line. When a video has several audio tracks, ffmpeg’s default stream is used — the first one — and there is no flag to choose another.

The list above is what the tests decode, not a claim about what ffmpeg can decode. A format outside it is not refused — it is simply unverified here.

Audio is read one window at a time rather than loaded whole, so peak memory does not grow with the length of the file: one hour of 16 kHz mono audio is 219 MiB held as a single array, and it is never held. Measured peak RSS of the tool on one hour of audio, before and after: 700 MiB → 479 MiB. The saving is in proportion to the duration, so short files gain nothing.

Convert it yourself if you prefer:

ffmpeg -i input.mp3 -ar 16000 -ac 1 output.wav

This site uses Just the Docs, a documentation theme for Jekyll.