You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

通用WAV文件摩尔斯码解码:音频与静音区分算法优化需求

解码含摩尔斯码的WAV文件:通用音频/静音区分算法需求

我正在编写Python程序解码含摩尔斯码的WAV文件,核心是将音频分割为音频块与静音块并确定音频平均时长,以此完成解码。

过往实现与局限

基于wave模块的实现

仅对程序生成的正弦波WAV文件有效,对自然声音的识别结果混乱,代码如下:

def read_audio_data(filename):
    # Returns data of a wav file as bytes.
    with wave.open(filename, 'rb') as wav:
        (nchannels, sampwidth, framerate, nframes, comptype, compname) = wav.getparams()
        frames = wav.readframes(nframes * nchannels)
        return frames


def frames_to_ints(wav_bytes):
    # Converts wav frames as bytes to a list of integers.
    return list(wav_bytes)


def find_mean_amplitude(frames):
    return mean(frames)


def rolling_average_ms(samples, sample_rate=44100):
    # The window represents an number of frames. For the standard sample rate of 44100, 1 millisecond
    # takes up 44.1 frames.
    moving_window = []
    rolling_average = []
    for i in samples[::2]:
        if len(moving_window) > sample_rate // 1000:
            del(moving_window[0])
        moving_window.append(i)
        rolling_average.append(mean(moving_window))
    return rolling_average

def sound_or_no_sound(samples):
    rolling_average = rolling_average_ms(samples)
    mean_amplitude = find_mean_amplitude(samples)

    on_or_off = []

    for i in rolling_average:
        if i < mean_amplitude:
            on_or_off.append(False)
        else:
            on_or_off.append(True)

    return on_or_off

基于scipy.wavfile的实现

可识别我录制的口哨摩尔斯码音频,但无法识别哼唱或计算机生成的音频文件,代码如下:

data = wavfile.read('blah_blah_blah.wav')

def split_into_ms(data):
    # Roughly groups the data by milliseconds.
    sample_rate = data[0]
    ms_length = sample_rate // 1000  # Roughly 1 millisecond
    values = data[1]
    ms = []
    for x in range(len(values))[ms_length - 1::ms_length]:
        ms.append(values[x - 44: x])
    return ms


def distinguish_sound_from_silence(ms, data):
    # Takes the milliseconds and decides if the sound is on or off. Adds "True" for on and "False" for off.
    data_stdev = np.std(data[1])
    output = []
    for x in ms:
        if np.std(x) >= data_stdev:
            output.append(True)
        else:
            output.append(False)
    return output

def group_by_sound_or_silence(bool_data):
    # Will return a list of blocks of sound and silence.
    output = []
    block = []

    for x in range(1, len(bool_data)):
        if bool_data[x] == bool_data[x - 1]:
            block.append(bool_data[x])
        else:
            output.append(block)
            block = []

    return output


def mean_block_length(blocks):
    return mean([len(block) for block in blocks])


def filter_short_blocks(blocks, tolerance=3):
    # Removes unusually short blocks.
    # Blocks of shorter length than the mean block length divided by the tolerance are removed.
    # Lower tolerance will filter more blocks.
    output = []
    mean_block_len = mean_block_length(blocks)
    for block in blocks:
        if len(block) > mean_block_len / tolerance:
            output.append(block)
    return output

需求

原本打算为自己的语音编写专用算法,但意识到针对每种声音单独适配的方式极不现实。现寻求一种通用算法,满足:

  • 可处理任意合理频率、任意波形的音频
  • 最好支持任意采样率(可依赖摩尔斯码频率相对稳定的特性)
  • 能准确区分音频与静音,进而完成摩尔斯码解码

内容的提问来源于stack exchange,提问作者HorrowShow

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.18 23:35:20