You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Sounddevice同步音频录播自动增益缩放问题求助

问题:Sounddevice库playrec接口自动缩放信号的解决方案?

我在树莓派上搭配HiFiBerry DAC + ADC PRO实现白噪声生成、传输,随后通过环回配置捕获并回放该信号。为降低延迟,采用Sounddevice库的playrec接口,但遇到的问题是发送的信号会被自动缩放到0-1区间。我原以为问题出在数据类型仅支持[-1,1],但已将dtype设置为np.int32,问题仍未解决。是否存在无需手动缩放数据的解决办法?

主程序文件

import Audio_Utils_Sound as audio
import numpy as np
import matplotlib.pyplot as plt
# This program is made to test the frequency spectre of a device

# Settings
sample_rate = 44100  # Sample rate for audio playback
duration = 3  # Duration of each test signal in seconds
amplitude = 1  # Amplitude of test signal
channels = 1  # 1 = mono, 2 = stereo
device = 0  # You can specify the ID or the name
dtype = np.int32  # Select the bit depth: np.int8, np.int16, np.int32(24bit doesn't exist so use np.int32)
num_bins = 512

audio.device_setup(channels, sample_rate, device)

num_samples = int(duration * sample_rate)
freqs = np.fft.fftfreq(num_samples, sample_rate)
noise = np.random.randn(num_samples)

recorded = audio.write_and_record(noise, sample_rate, channels, dtype)  # Record signal from device
print(max(noise))
print(max(recorded))

# Calculate FFT and decibels of input
amplitude_db_input, freqs_input = audio.fft(noise, sample_rate, duration)
amplitude_db_output, freqs_output = audio.fft(recorded, sample_rate, duration)
amplitude_db_output_gain, freqs_output_gain = audio.fft_gain(noise, recorded, sample_rate, duration)

# Plot frequency response graphs
plt.figure(figsize=(16, 8))

plt.subplot(2, 2, 1)
plt.plot(freqs_input, amplitude_db_input, linestyle='-')
plt.xscale('log')
plt.xlabel('Frequency (Hz)')
plt.ylabel('Amplitude (dB)')
plt.title('Input Noise')
plt.grid(True)

plt.subplot(2, 2, 2)
plt.plot(freqs_output, amplitude_db_output, linestyle='-')
plt.xscale('log')
plt.xlabel('Frequency (Hz)')
plt.ylabel('Amplitude (dB)')
plt.title('Frequency Response')
plt.grid(True)

plt.subplot(2, 2, 3)
plt.plot(freqs_output_gain, amplitude_db_output_gain, linestyle='-')
plt.xscale('log')
plt.xlabel('Frequency (Hz)')
plt.ylabel('Gain (dB)')
plt.title('Frequency Response')
plt.grid(True)

plt.tight_layout()
plt.show()

工具类文件

import numpy as np
import sounddevice as sd

def fft(signal, sample_rate, duration):
    num_samples = int(duration * sample_rate)
    
    fft_output = np.fft.fft(signal)
    amplitude_output = abs(fft_output)
    amplitude_db_output = 20 * np.log10(amplitude_output)          
    freqs_output = np.fft.fftfreq(num_samples, sample_rate) # Using np.fft.fftfreq negative frequency's are also created, so we filter this out
    return amplitude_db_output[freqs_output >= 0], freqs_output[freqs_output >= 0]

def fft_gain(output_signal, input_signal, sample_rate, duration):
    num_samples = int(duration * sample_rate)
    
    fft_output = np.fft.fft(output_signal)
    amplitude_output = np.abs(fft_output)
    
    fft_input = np.fft.fft(input_signal)
    amplitude_input = np.abs(fft_input)
    
    gain_db_output = 20 * np.log10(amplitude_input / amplitude_output)
    
    freqs_output = np.fft.fftfreq(num_samples, sample_rate) # Using np.fft.fftfreq negative frequency's are also created, so we filter this out
    
    return gain_db_output[freqs_output >= 0], freqs_output[freqs_output >= 0]

def write_and_record(data, sample_rate, channels, dtype):
    # Play audio and record simultaneously
    frames = sd.playrec(data, samplerate=sample_rate, channels=channels, blocking=True)
    sd.wait()
    return frames[:, 0]     # Returns(frames, channels)

def device_setup(channels, sample_rate, device):
    sd.default.channels = channels
    sd.default.samplerate = sample_rate
    sd.default.device = device
    print('Devices:\n' + str(sd.query_devices()) + '\n')  # print available devices
    # print('Host APIs:\n' + str(sd.query_hostapis())+ '\n')     # print available host APIs
    print('Input:\n' + str(sd.check_input_settings()) + '\n')  # print available devices
    print('Output:\n' + str(sd.check_output_settings()) + '\n\n')  # print available devices

我最初以为是HiFiBerry的问题,但使用PyAudio时不存在此缩放问题。当我发送最大值为4.52的信号时,返回的最大值仍为0.9999,我需要计算增益,因此希望禁用该库引入的自动增益缩放功能。我在Sounddevice文档中未找到相关内容,是否有人遇到过此问题?


内容的提问来源于stack exchange,提问作者Gilles De Roo

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.26 11:16:29