如何使用FFmpeg检测视频文件中的静音音频通道/轨道
如何使用FFmpeg检测视频文件中的静音音频通道/轨道
我需要能够扫描视频文件,在事先不知道文件音频映射的情况下,报告哪些音频轨道/通道是静音的。我在一个包含16个音频通道的文件上尝试了这个FFmpeg(v4.1)命令:
$ ffmpeg -i foo5.mxf -map 0:a -af astats -f null - ... Stream mapping: Stream #0:1 -> #0:0 (pcm_s24be (native) -> pcm_s16le (native)) Stream #0:2 -> #0:1 (pcm_s24be (native) -> pcm_s16le (native)) Stream #0:3 -> #0:2 (pcm_s24be (native) -> pcm_s16le (native)) Stream #0:4 -> #0:3 (pcm_s24be (native) -> pcm_s16le (native)) Stream #0:5 -> #0:4 (pcm_s24be (native) -> pcm_s16le (native)) Stream #0:6 -> #0:5 (pcm_s24be (native) -> pcm_s16le (native)) Stream #0:7 -> #0:6 (pcm_s24be (native) -> pcm_s16le (native)) Stream #0:8 -> #0:7 (pcm_s24be (native) -> pcm_s16le (native)) Stream #0:9 -> #0:8 (pcm_s24be (native) -> pcm_s16le (native)) Stream #0:10 -> #0:9 (pcm_s24be (native) -> pcm_s16le (native)) Stream #0:11 -> #0:10 (pcm_s24be (native) -> pcm_s16le (native)) Stream #0:12 -> #0:11 (pcm_s24be (native) -> pcm_s16le (native)) Stream #0:13 -> #0:12 (pcm_s24be (native) -> pcm_s16le (native)) Stream #0:14 -> #0:13 (pcm_s24be (native) -> pcm_s16le (native)) Stream #0:15 -> #0:14 (pcm_s24be (native) -> pcm_s16le (native)) Stream #0:16 -> #0:15 (pcm_s24be (native) -> pcm_s16le (native)) ... [Parsed_astats_0 @ 0x9123c0] Channel: 1 [Parsed_astats_0 @ 0x9123c0] DC offset: 0.000070 [Parsed_astats_0 @ 0x9123c0] Min level: -170426368.000000 [Parsed_astats_0 @ 0x9123c0] Max level: 172195840.000000 [Parsed_astats_0 @ 0x9123c0] Min difference: 0.000000 [Parsed_astats_0 @ 0x9123c0] Max difference: 58281984.000000 [Parsed_astats_0 @ 0x9123c0] Mean difference: 3089081.855667 [Parsed_astats_0 @ 0x9123c0] RMS difference: 4825348.720329 [Parsed_astats_0 @ 0x9123c0] Peak level dB: -21.918144 [Parsed_astats_0 @ 0x9123c0] RMS level dB: -34.035425 [Parsed_astats_0 @ 0x9123c0] RMS peak dB: -27.320939 [Parsed_astats_0 @ 0x9123c0] RMS trough dB: -48.414915 [Parsed_astats_0 @ 0x9123c0] Crest factor: 4.035190 [Parsed_astats_0 @ 0x9123c0] Flat factor: 0.000000 [Parsed_astats_0 @ 0x9123c0] Peak count: 2 [Parsed_astats_0 @ 0x9123c0] Bit depth: 20/20 [Parsed_astats_0 @ 0x9123c0] Dynamic range: 98.493854 [Parsed_astats_0 @ 0x9123c0] Zero crossings: 10216 [Parsed_astats_0 @ 0x9123c0] Zero crossings rate: 0.042383 [Parsed_astats_0 @ 0x9123c0] Overall [Parsed_astats_0 @ 0x9123c0] DC offset: 0.000070 [Parsed_astats_0 @ 0x9123c0] Min level: -170426368.000000 [Parsed_astats_0 @ 0x9123c0] Max level: 172195840.000000 [Parsed_astats_0 @ 0x9123c0] Min difference: 0.000000 [Parsed_astats_0 @ 0x9123c0] Max difference: 58281984.000000 [Parsed_astats_0 @ 0x9123c0] Mean difference: 3089081.855667 [Parsed_astats_0 @ 0x9123c0] RMS difference: 4825348.720329 [Parsed_astats_0 @ 0x9123c0] Peak level dB: -21.918144 [Parsed_astats_0 @ 0x9123c0] RMS level dB: -34.035425 [Parsed_astats_0 @ 0x9123c0] RMS peak dB: -27.320939 [Parsed_astats_0 @ 0x9123c0] RMS trough dB: -48.414915 [Parsed_astats_0 @ 0x9123c0] Flat factor: 0.000000 [Parsed_astats_0 @ 0x9123c0] Peak count: 2.000000 [Parsed_astats_0 @ 0x9123c0] Bit depth: 20/20 [Parsed_astats_0 @ 0x9123c0] Number of samples: 241040 [Parsed_astats_0 @ 0x4def40] Channel: 1 [Parsed_astats_0 @ 0x4def40] DC offset: 0.000070 [Parsed_astats_0 @ 0x4def40] Min level: -170426368.000000 [Parsed_astats_0 @ 0x4def40] Max level: 172199936.000000 [Parsed_astats_0 @ 0x4def40] Min difference: 0.000000 [Parsed_astats_0 @ 0x4def40] Max difference: 58277888.000000 [Parsed_astats_0 @ 0x4def40] Mean difference: 3089075.891088 [Parsed_astats_0 @ 0x4def40] RMS difference: 4825346.346176 [Parsed_astats_0 @ 0x4def40] Peak level dB: -21.917938 [Parsed_astats_0 @ 0x4def40] RMS level dB: -34.035425 [Parsed_astats_0 @ 0x4def40] RMS peak dB: -27.320942 [Parsed_astats_0 @ 0x4def40] RMS trough dB: -48.414908 [Parsed_astats_0 @ 0x4def40] Crest factor: 4.035287 [Parsed_astats_0 @ 0x4def40] Flat factor: 0.000000 [Parsed_astats_0 @ 0x4def40] Peak count: 2 [Parsed_astats_0 @ 0x4def40] Bit depth: 20/20 [Parsed_astats_0 @ 0x4def40] Dynamic range: 98.494061 [Parsed_astats_0 @ 0x4def40] Zero crossings: 10220 [Parsed_astats_0 @ 0x4def40] Zero crossings rate: 0.042400 [Parsed_astats_0 @ 0x4def40] Overall [Parsed_astats_0 @ 0x4def40] DC offset: 0.000070 [Parsed_astats_0 @ 0x4def40] Min level: -170426368.000000 [Parsed_astats_0 @ 0x4def40] Max level: 172199936.000000 [Parsed_astats_0 @ 0x4def40] Min difference: 0.000000 [Parsed_astats_0 @ 0x4def40] Max difference: 58277888.000000 [Parsed_astats_0 @ 0x4def40] Mean difference: 3089075.891088 [Parsed_astats_0 @ 0x4def40] RMS difference: 4825346.346176 [Parsed_astats_0 @ 0x4def40] Peak level dB: -21.917938 [Parsed_astats_0 @ 0x4def40] RMS level dB: -34.035425 [Parsed_astats_0 @ 0x4def40] RMS peak dB: -27.320942 [Parsed_astats_0 @ 0x4def40] RMS trough dB: -48.414908 [Parsed_astats_0 @ 0x4def40] Flat factor: 0.000000 [Parsed_astats_0 @ 0x4def40] Peak count: 2.000000 [Parsed_astats_0 @ 0x4def40] Bit depth: 20/20 [Parsed_astats_0 @ 0x4def40] Number of samples: 241040 ...
通过管道传给grep(再加一点过滤),我可以汇总这些信息并检查音频通道的峰值电平:
$ ffmpeg -i foo5.mxf -map 0:a -af astats -f null - 2>&1 | grep -E "(Channel:|Peak level|Overall)" [Parsed_astats_0 @ 0x9123c0] Channel: 1 [Parsed_astats_0 @ 0x9123c0] Peak level dB: -21.918144 [Parsed_astats_0 @ 0x9123c0] Overall [Parsed_astats_0 @ 0x9123c0] Peak level dB: -21.918144 [Parsed_astats_0 @ 0x4def40] Channel: 1 [Parsed_astats_0 @ 0x4def40] Peak level dB: -21.917938 [Parsed_astats_0 @ 0x4def40] Overall [Parsed_astats_0 @ 0x4def40] Peak level dB: -21.917938 [Parsed_astats_0 @ 0xb92180] Channel: 1 [Parsed_astats_0 @ 0xb92180] Peak level dB: -6153.053111 [Parsed_astats_0 @ 0xb92180] Overall [Parsed_astats_0 @ 0xb92180] Peak level dB: -6153.053111 ...
这样如果我过滤掉峰值电平小于-120dB的结果,基本能得到我需要的信息。但问题是它把所有通道都标识为“Channel 1”——可能是因为每个音频流本身只有一个通道。但我在原始输出里找不到任何能把每个“音频统计区段”(也就是包含从“Channel:”到“Zero crossings rate”的部分)和对应的流(比如0:0、0:1等)关联起来的信息。我能不能只根据输出的顺序来可靠地推断这种对应关系?
或者,我是不是应该使用不同的滤镜(而不是astats)或者参数设置来获取这些信息?
备注:内容来源于stack exchange,提问作者kfank
相关产品推荐
相关产品推荐

