如何修复Python读取hddtemp输出时的'utf-8'解码错误?
问题描述
我编写了一段Python3代码用于读取硬盘温度,代码如下:
device = '/dev/%s' % self.device tempRe = re.compile('%s:.*:(.*)' % device) hddtemp = os.popen('/usr/sbin/hddtemp -q %s' % device) for line in hddtemp: temp = re.findall(tempRe, line) if temp: self['temp'].setText(_('Disk temperature: %s') % temp[0].lstrip()) hddtemp.close()
部分硬盘可正常运行,例如输出:
/dev/sda: TOSHIBA MK2565GSX HR: 34°C
但部分硬盘的输出包含特殊字符,例如:
/dev/sda: Thinklife SSD ST600 240G ▒: 23°C
此时代码会抛出以下错误:
for line in hddtemp: File "<frozen codecs>", line 322, in decode UnicodeDecodeError: 'utf-8' codec can't decode byte 0x80 in position 51: invalid start byte
请问该如何修复此问题?
修复方案
方法1:指定编码并处理解码错误
os.popen默认使用系统编码解码输出,但hddtemp的输出可能并非UTF-8,导致解码失败。推荐改用subprocess模块,更灵活控制编码和错误处理:
import subprocess import re device_path = f'/dev/{self.device}' cmd = ['/usr/sbin/hddtemp', '-q', device_path] # 用errors='replace'替换无法解码的字符,不影响后续匹配 output = subprocess.check_output(cmd, errors='replace').decode('utf-8') # 若已知hddtemp输出编码(如latin-1),可直接指定:output = subprocess.check_output(cmd).decode('latin-1') # 用re.escape避免设备路径中的特殊字符干扰正则匹配 temp_re = re.compile(f'{re.escape(device_path)}:.*:(.*)') match_result = temp_re.search(output) if match_result: temp_value = match_result.group(1).lstrip() self['temp'].setText(_('Disk temperature: %s') % temp_value)
方法2:二进制模式读取后处理
不确定编码时,可先以二进制模式读取输出,再处理解码:
import os import re device = '/dev/%s' % self.device # 生成二进制正则表达式 tempRe = re.compile(b'%s:.*:(.*)' % device.encode('utf-8')) # 以二进制模式打开管道 hddtemp = os.popen('/usr/sbin/hddtemp -q %s' % device, mode='rb') for line in hddtemp: temp = re.findall(tempRe, line) if temp: # 解码时忽略无法处理的字符 temp_str = temp[0].lstrip().decode('utf-8', errors='ignore') self['temp'].setText(_('Disk temperature: %s') % temp_str) hddtemp.close()
额外优化建议
- 优先使用
subprocess模块替代os.popen,前者是官方推荐的子进程处理方式,功能更全面且安全性更高 - 正则匹配时加入
re.escape(device_path),避免设备路径中的特殊符号破坏正则规则
内容的提问来源于stack exchange,提问作者Fair Bird
相关产品推荐
相关产品推荐

