使用AudioSegment.from_file()触发AttributeError: 'NoneType'无'read'属性
解决语音转文本项目中的AttributeError问题
错误原因
你遇到的AttributeError: 'NoneType' object has no attribute 'read',核心问题是错误调用了AudioSegment的类方法:
你写的是AudioSegment().from_file(data),先实例化了一个空的AudioSegment对象,再去调用类方法from_file,这种用法不符合pydub的API规范,导致内部处理时传入了无效的None值,触发了读取操作的错误。
from_file是AudioSegment的类方法,应该直接通过类本身调用,不需要先创建实例。
解决方案
将代码中的:
clip = AudioSegment().from_file(data)
修改为:
clip = AudioSegment.from_file(data)
另外,为了避免因无有效音频输入导致的潜在问题,可以给r.listen(source)添加超时或能量阈值判断,防止程序一直等待输入或捕捉到无效音频:
audio = r.listen(source, timeout=5, phrase_time_limit=10)
修正后的完整代码
import torch from transformers import Wav2Vec2ForCTC, Wav2Vec2Processor import speech_recognition as sr from pydub import AudioSegment import io r = sr.Recognizer() tokenizer = Wav2Vec2Processor.from_pretrained("cahya/wav2vec2-base-turkish") model = Wav2Vec2ForCTC.from_pretrained("cahya/wav2vec2-base-turkish") with sr.Microphone(sample_rate=16000) as source: print("You can start speaking now") while True: try: audio = r.listen(source, timeout=5, phrase_time_limit=10) data = io.BytesIO(audio.get_wav_data()) clip = AudioSegment.from_file(data) x = torch.FloatTensor(clip.get_array_of_samples()) inputs = tokenizer(x, sampling_rate = 16000, return_tensors = 'pt', padding = 'longest' ).input_values logits = model(inputs).logits tokens = torch.argmax(logits, dim = -1) text = tokenizer.batch_decode(tokens) print(str(text).lower()) except sr.WaitTimeoutError: print("No audio detected, please speak again")
内容的提问来源于stack exchange,提问作者Awrelo5
相关产品推荐
相关产品推荐

