You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将FFmpeg音频流管道至Sox Rec,实现Icecast直播音频分段录音?

Icecast Stream to Sox: Auto-Split Recordings on Silence

Got it, let's tackle this directly. Here's a fully functional command chain that pulls your Icecast stream, pipes it to Sox, and automatically splits recordings into segments only when there's active speech—stopping on silence and starting a new file when talking resumes:

ffmpeg -i http://icecast.radiovox.org:8000/live.ogg -vn -acodec pcm_s16le -ar 32000 -ac 1 -f wav - | rec - rate 32k silence 1 0.1 3% 1 3.0 3% "talkshow_segment_%03d.mp3"

Let's break down each part to understand how it works, plus how to tweak it for your specific needs:

FFmpeg: Stream Fetching & Format Prep

FFmpeg handles pulling the stream and converting it to a format Sox can process without hiccups:

  • -i http://icecast.radiovox.org:8000/live.ogg: Target Icecast stream URL
  • -vn: Disables video processing (unnecessary for an audio-only stream)
  • -acodec pcm_s16le: Converts the stream to 16-bit PCM audio—Sox natively supports this, avoiding format compatibility issues
  • -ar 32000: Sets the sample rate to 32kHz, matching Sox's rate 32k to skip redundant resampling
  • -ac 1: Converts to mono audio (ideal for talk shows, and cuts down on file size)
  • -f wav -: Outputs the converted audio as WAV to standard output (-), which feeds directly into Sox via the pipe

Sox: Silence Detection & Auto-Split Logic

The rec command takes the piped audio and handles the intelligent recording rules:

  • rec -: Reads input from standard output (FFmpeg's stream)
  • rate 32k: Confirms the sample rate (matches FFmpeg's output for efficiency)
  • silence 1 0.1 3% 1 3.0 3%: This is the core silence detection rule—here's what each part means:
    • 1 0.1 3%: Stop recording when 1 segment of silence (volume below 3% of max) lasts at least 0.1 seconds
    • 1 3.0 3%: Start a new recording file when 1 segment of speech (volume above 3% of max) lasts at least 3.0 seconds
  • "talkshow_segment_%03d.mp3": Output file template—%03d generates sequential filenames like talkshow_segment_001.mp3, talkshow_segment_002.mp3, so you never overwrite existing files

Quick Tweaks for Your Use Case

If the split logic isn't behaving exactly as you want, adjust these parameters:

  • Silence Sensitivity: Change the 3% values to tweak how strict the silence detection is. Use 5% if you want to cut off quieter background noise, or 2% if softer speech is being mistaken for silence.
  • Pause Tolerance: Modify 0.1 (stop delay) if short pauses in the show are triggering unwanted splits—try 0.5 to tolerate longer gaps. Adjust 3.0 (start delay) if you want to start recording faster when speech resumes (e.g., 1.0 instead).
  • Output Format: Swap .mp3 for .wav or .flac if you need lossless recordings—Sox will handle the format conversion automatically.

内容的提问来源于stack exchange,提问作者Red-Tune-84

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.26 10:52:53