You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于FFmpeg的转码音频检测工具技术方案问询

问题描述

我开发了一款自用CLI工具,用于检测转码音频文件,目标识别由低质量源转码而来的“伪”高码率音频。因信号处理非我专长,特寻求以下技术反馈:

  1. 该方案是否合理,分析过程存在明显错误吗?
  2. 有无更高效的实现方式(不采用并行化)?
  3. 该方案存在哪些未覆盖的边缘场景?
# Analyzes the frequency content of audio files to verify whether
# the reported bitrate reflects the actual audio quality, or whether
# the file has been transcoded from a lower-quality source.

import os
import time
import subprocess
import numpy as np

LIBS_PATH = "../libs"

def generate_spectrogram(path: str):
    base        = os.path.splitext(os.path.basename(path))[0]
    output_path = f"{base}_spectrogram.png"

    subprocess.run(
        [
            f"{LIBS_PATH}/ffmpeg",
            "-v",     "error",                                                        # Suppress all non-error messages
            "-i",     path,                                                           # Input file
            "-y",     output_path,                                                    # Output file, overwrite if exists
            "-lavfi", "showspectrumpic=s=400x250:legend=1:scale=log:color=intensity", # Settings for the spectrogram
        ],
        check = True,
    )

def read_audio_chunk(path: str, start_at_sec: float):
    DURATION_SEC = 2

    # Use `ffmpeg` to decode the chunk of the audio file to raw PCM samples.
    result = subprocess.run(
        [
            f"{LIBS_PATH}/ffmpeg",
            "-v",  "error",           # Suppress all non-error messages
            "-ss", str(start_at_sec), # Seek to this position in the track before decoding
            "-i",  path,              # Input file
            "-t",  str(DURATION_SEC), # Decode only this many seconds
            "-ac", "1",               # Downmix to mono to simplify the analysis
            "-f",  "f32le",           # Set output as raw little-endian float32 samples with no container
            "-"                       # Write to stdout
        ],
        capture_output = True,
    )

    return np.frombuffer(result.stdout, dtype = np.float32)

# Finds the frequency cutoff of an audio file.
#
# Scan the PCM samples in 4 different chunks spaced through the track. For each chunk,
# apply Real Fast Fourier Transform (RFFT) to separate the frequencies into bins, and analyze
# them to find the highest frequency bin across the track that is not "silence".
def find_frequency_cutoff(path: str):
    # Use `ffprobe` to get the metadata we need to interpret the PCM samples
    ffprobe_out = subprocess.run(
        [
            f"{LIBS_PATH}/ffprobe",
            "-v",              "error",                                           # Suppress all non-error messages
            "-select_streams", "a:0",                                             # Target only the first audio stream, ignoring video streams
            "-show_entries",   "stream=sample_rate,bit_rate,codec_name,duration", # Fields to extract, one per output line
            "-of",             "default=noprint_wrappers=1:nokey=1",              # Output values only, no keys or section headers
            path                                                                  # Input file
        ],
        capture_output = True,
        text           = True,
    )

    lines        = ffprobe_out.stdout.strip().splitlines()
    codec        = lines[0]
    sample_rate  = int(lines[1])
    duration_sec = float(lines[2])
    bitrate_bps  = int(lines[3])

    # =====

    max_cutoff_hz = -1.0

    for pct in [0.25, 0.4, 0.6, 0.75]:
        start_at_sec = duration_sec * pct
        chunk        = read_audio_chunk(path, start_at_sec)

        # Apply a Hanning window to reduce spectral leakage, then compute
        # the frequency spectrum using RFFT.
        chunk    = chunk * np.hanning(len(chunk))
        spectrum = np.abs(np.fft.rfft(chunk))

        peak = spectrum.max()

        if peak == 0:
            continue  # This chunk is silent

        # Normalize to `[0, 1]` so all files share the same reference level.
        spectrum = spectrum / peak

        # Convert to dBFS (decibels relative to full scale)
        # 0 dBFS is the loudest bin; everything else is negative.
        spectrum = 20 * np.log10(spectrum)

        # =====

        bin_resolution_hz = sample_rate / len(chunk) # Hz per bin

        BAND_SIZE = 50

        # Scan from the highest frequency bin downward, looking for the
        # first `band` where most bins are above the noise floor.
        for i in range(len(spectrum), BAND_SIZE, -1):
            band = spectrum[i - BAND_SIZE : i]

            # Sounds below this dBFS level are considered "silence".
            NOISE_FLOOR_DBFS = -100

            # If most bins are above the noise floor, it's probably real content 
            # rather than just noise. If so, the upper edge of this band is the cutoff.
            if (band > NOISE_FLOOR_DBFS).mean() >= 0.5:
                cutoff_hz     = (i - 1) * bin_resolution_hz
                max_cutoff_hz = max(max_cutoff_hz, cutoff_hz)
                break

    # If we never found a band above the noise floor, the file is likely silent/corrupt.
    if max_cutoff_hz < 0:
        max_cutoff_hz = None

    return path, codec, bitrate_bps, max_cutoff_hz

# =====

FILES = [
    os.path.join("../downloads", f)
    for f in os.listdir("../downloads")
    if f.endswith(".mp3")
]

COL_SIZE = [20, 6, 10, 10]
header   = f"{'File':<{COL_SIZE[0]}} {'Codec':<{COL_SIZE[1]}} {'Bitrate':>{COL_SIZE[2]}} {'Cutoff':>{COL_SIZE[3]}}"

print(header)
print("-" * len(header))

started_at = time.time()

for p in FILES:
    path, codec, bitrate, cutoff_hz = find_frequency_cutoff(p)

    if cutoff_hz is None:
        cutoff_str = "N/A"
    else:
        cutoff_str = f"{cutoff_hz:.0f} Hz"

    bitrate = f"{bitrate / 1000:.0f} kBPS"
    name    = os.path.basename(p)

    if len(name) > COL_SIZE[0]:
        name = name[:COL_SIZE[0] - 4] + "..."

    print(f"{name:<{COL_SIZE[0]}} {codec:<{COL_SIZE[1]}} {bitrate:>{COL_SIZE[2]}} {cutoff_str:>{COL_SIZE[3]}}")

elapsed_sec = time.time() - started_at
print(f"\nAnalyzed {len(FILES)} files in {elapsed_sec:.1f} seconds")

(代码中的大量注释是留给未来自己的,用于回忆实现逻辑)


技术反馈

1. 方案合理性与潜在错误

整体思路没问题——低码率编码会丢弃高频信号,转成高码率后也无法恢复,通过检测高频 cutoff 识别伪高码率的逻辑成立。但有几个细节存在隐患:

  • 固定噪声阈值不合理:NOISE_FLOOR_DBFS = -100是硬编码值,不同音频的噪声基底差异极大(比如老录音、现场录音的噪声更高),固定阈值会导致误判。建议改成自适应计算当前chunk的噪声基底,比如取频谱中最低20%的均值再加10dB作为阈值。
  • 峰值归一化的缺陷:用频谱峰值做归一化后转dBFS,会把噪声也放大到接近0dB,导致噪声频段被误判为有效信号。应该改用**RMS(均方根)**或者基于整体能量的归一化方式,避免噪声被过度放大。
  • ffprobe输出解析依赖顺序:直接按lines[0]到lines[3]提取字段,完全依赖ffprobe的输出顺序,但实际输出顺序不一定严格匹配你指定的stream=sample_rate,bit_rate,codec_name,duration。最好改用-of json格式输出,解析JSON获取对应字段,可靠性更高。
  • FFT分辨率冗余:2秒的chunk做FFT,频率分辨率过高(比如44.1kHz采样下分辨率约0.5Hz),既增加计算量,也对cutoff判断没有帮助。可以缩短chunk时长到1秒,或者分帧做STFT后取平均,结果更稳定,计算也更快。

2. 非并行化的高效实现方式

  • 复用解码流程:现在每个chunk都单独调用ffmpeg解码,进程启动的开销极大。改成一次性解码整个音频到内存中的PCM数据,再在内存中切分chunk,避免多次启动子进程的开销。
  • 改用STFT分析:不需要对每个chunk做完整RFFT,用短时傅里叶变换(STFT)分帧处理,取各帧频谱的平均值,结果更稳定,计算效率也更高。
  • 简化频谱计算精度:频谱计算用float32足够,不需要更高精度;cutoff判断可以放宽到100Hz的粒度,不用精确到1Hz,减少计算量。
  • 提前过滤无效文件:处理前先检查文件的有效性(比如用ffprobe快速判断是否是可解码的音频),避免对损坏文件做无用计算。
  • 减少重复操作:比如generate_spectrogram函数如果不是必须的,可以考虑默认不生成,或者按需触发,减少不必要的ffmpeg调用。

3. 未覆盖的边缘场景

  • 自然低高频的音频:比如人声独白、老唱片、低保真风格音乐,它们的自然高频cutoff本来就很低,会被误判为伪高码率。需要结合codec和标称bitrate做规则判断:比如MP3 320kbps的正常cutoff应该接近20kHz,如果标称320kbps但cutoff只有16kHz,才判定为伪高码率;而本身就是低cutoff的高码率文件则放过。
  • 带高频噪声的伪高码率文件:有些低码率转高码率时会添加随机高频噪声伪装,当前逻辑会把噪声当成有效信号,无法识别。需要区分有效信号和噪声:有效信号的高频段频谱是连续的,而噪声是随机离散的,可以通过计算频段的方差来判断。
  • 多声道音频:现在强制转成单声道,但有些多声道音频(比如5.1)的高频可能只存在于某个声道,转单声道后会被平均掉,导致cutoff判断偏低。应该单独分析每个声道的高频,取最大值作为最终cutoff。
  • 可变比特率(VBR)文件:VBR文件的不同段落bitrate差异大,有些段落可能是高码率,有些是低码率,只取4个chunk可能漏判。需要增加采样chunk的数量,或者重点检测bitrate波动大的段落。
  • 无损转码的伪无损文件:比如把低码率MP3转成FLAC,当前工具会检测到cutoff低,但FLAC是无损格式,这时候需要结合codec判断:无损格式的cutoff应该接近采样率的一半(奈奎斯特频率),如果远低于这个值,才判定为伪无损。
  • 短音频文件:时长小于2秒的文件,当前chunk采样逻辑会出现start_at_sec超过文件时长的情况,导致读取失败。需要单独处理短文件,直接分析整个文件的频谱。

内容的提问来源于stack exchange,提问作者petdomaa100

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.01 22:17:26