You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用AudioSegment.from_file()触发AttributeError: 'NoneType'无'read'属性

解决语音转文本项目中的AttributeError问题

错误原因

你遇到的AttributeError: 'NoneType' object has no attribute 'read',核心问题是错误调用了AudioSegment的类方法:
你写的是AudioSegment().from_file(data),先实例化了一个空的AudioSegment对象,再去调用类方法from_file,这种用法不符合pydub的API规范,导致内部处理时传入了无效的None值,触发了读取操作的错误。

from_file是AudioSegment的类方法,应该直接通过类本身调用,不需要先创建实例。

解决方案

将代码中的:

clip = AudioSegment().from_file(data)

修改为:

clip = AudioSegment.from_file(data)

另外,为了避免因无有效音频输入导致的潜在问题,可以给r.listen(source)添加超时或能量阈值判断,防止程序一直等待输入或捕捉到无效音频:

audio = r.listen(source, timeout=5, phrase_time_limit=10)

修正后的完整代码

import torch
from transformers import Wav2Vec2ForCTC, Wav2Vec2Processor
import speech_recognition as sr
from pydub import AudioSegment
import io

r = sr.Recognizer()

tokenizer = Wav2Vec2Processor.from_pretrained("cahya/wav2vec2-base-turkish")
model = Wav2Vec2ForCTC.from_pretrained("cahya/wav2vec2-base-turkish")

with sr.Microphone(sample_rate=16000) as source: 
    print("You can start speaking now") 
    while True: 
        try:
            audio = r.listen(source, timeout=5, phrase_time_limit=10) 
            data = io.BytesIO(audio.get_wav_data())
            
            clip = AudioSegment.from_file(data)
            x    = torch.FloatTensor(clip.get_array_of_samples())
            
            inputs = tokenizer(x, sampling_rate = 16000, return_tensors = 'pt', padding = 'longest' ).input_values
            logits = model(inputs).logits
            tokens = torch.argmax(logits, dim = -1)
            text   = tokenizer.batch_decode(tokens)
            
            print(str(text).lower())
        except sr.WaitTimeoutError:
            print("No audio detected, please speak again")

内容的提问来源于stack exchange,提问作者Awrelo5

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.05 12:02:17