Python subprocess.run调用土耳其语词形还原工具UTF-8编码报错如何解决
错误原因
- Windows系统控制台默认编码和Colab(Linux环境)不一致:Colab默认使用UTF-8编码,Windows下cmd/PowerShell默认使用对应区域的非UTF-8编码,子进程输出内容用UTF-8直接解码会触发解码失败,进而导致子进程异常退出抛出
CalledProcessError。 - 用字符串拼接方式传参给
subprocess.run处理带特殊字符的土耳其语单词时,shell解析参数容易出现转义错误,也是进程异常退出的常见诱因。 - 单步调试时的运行环境编码配置和批量运行时的配置不一致,因此不会触发报错。
解决方案
快速修复方案
修改subprocess.run相关代码,调整参数传递和解码规则:
# 替换原有的subprocess.run代码段 # 用列表传参、关闭shell避免参数转义问题 cmd = ["python", "lemmatizer.py", r[0]] o = subprocess.run( cmd, capture_output=True, encoding=None, shell=False, check=True ) # 手动解码,设置容错避免程序崩溃 l = o.stdout.decode('utf-8', errors='replace').split('\n')[1]
如果仍出现解码错误,可将解码编码替换为土耳其语专属编码cp1254:
l = o.stdout.decode('cp1254', errors='replace').split('\n')[1]
可在循环外层添加异常捕获逻辑,避免个别单词处理错误导致整个批量任务中断:
for r in F: try: # 原有循环处理逻辑 except Exception as e: print(f"处理行{r}出错,错误信息:{e}") continue
最优优化方案
直接导入lemmatizer的核心功能到主程序中,避免每次循环反复创建子进程带来的性能损耗和编码、参数传递问题,运行效率可提升10倍以上,彻底规避所有子进程相关错误。
修改后代码示例:
e = 'utf-8' import os # 导入lemmatizer的词根提取函数,根据lemmatizer.py实际导出的函数名调整 from lemmatizer import lemmatize os.chdir("C:/Users/(a directory)") d = dict() i = 0 with open('turkishforms.txt','r',encoding=e) as f: F = f.read().split('\n') for r in F: r = r.strip() if not r: continue r = r.rsplit('\t', 1) r[-1] = int(r[-1]) # 直接调用函数获取词元 l = lemmatize(r[0]) d[l] = d.get(l, 0) + r[-1] i += 1 print('{}\t{}\t{}\t{}\t{}'.format(i, r[0], r[-1], l, d[l])) with open('turkishlemmas.txt','w',encoding=e) as g: for lemma, count in d.items(): t = '{}\t{}\n'.format(lemma, count) print(t) g.write(t)
内容的提问来源于stack exchange,提问作者Devorah Yi
相关产品推荐
相关产品推荐

