Python3.6执行XML解析代码时遭遇UTF-8 UnicodeEncodeError问题排查
解决BNC语料库XML解析时的UnicodeEncodeError问题
你的代码逻辑没问题,但踩了两个Python里常见的编码和路径坑,导致触发了UnicodeEncodeError。我来帮你一步步修复:
问题1:写入文件时未指定UTF-8编码
你读取XML文件时明确指定了encoding="utf-8",但写入输出文件时没指定编码——在Windows系统上,Python默认会用GBK编码写入文件,而BNC语料里肯定存在GBK无法编码的特殊Unicode字符,这就直接引发了编码错误。
问题2:路径字符串的转义陷阱
你的路径"C:\Users\####\Desktop\BNC2\[A00-ZZZ]*.xml"里的\U会被Python当成Unicode转义序列,要么触发语法错误,要么导致程序找不到正确的文件路径。
修正后的完整代码
#!/usr/bin/env python3 import glob import xml.etree.ElementTree as ET # 用原始字符串避免路径转义问题(推荐写法) filenames = glob.glob(r"C:\Users\####\Desktop\BNC2\[A00-ZZZ]*.xml") out_lines = [] for filename in filenames: with open(filename, 'r', encoding="utf-8") as content: tree = ET.parse(content) root = tree.getroot() for w in root.iter('w'): lemma = w.get('hw') pos = w.get('pos') tag = w.get('c5') # 处理可能的None值,避免拼接时触发TypeError lemma = lemma if lemma is not None else "" pos = pos if pos is not None else "" tag = tag if tag is not None else "" # 用f-string让字符串拼接更简洁 out_lines.append(f"{w.text},{lemma},{pos},{tag}") # 写入时指定UTF-8编码,同时设置newline=''避免Windows下多出空行 with open(r"C:\Users\####\Desktop\bnc_parsed_output.csv", 'w', encoding="utf-8", newline='') as out_file: out_file.write('\n'.join(out_lines))
关键修改点说明
- 路径处理:在路径字符串前加
r变成原始字符串,让Python不把反斜杠解析为转义字符,避免路径识别错误。 - 写入编码:打开输出文件时明确指定
encoding="utf-8",确保所有Unicode字符都能正确写入文件。 - None值防护:新增对
lemma、pos、tag的None判断,防止XML节点缺失这些属性时程序崩溃。 - 容错优化:如果还是遇到极个别无法编码的特殊字符,可以在写入时加上
errors="replace"参数,把无法编码的字符替换为?,保证程序正常运行:with open(r"C:\Users\####\Desktop\bnc_parsed_output.csv", 'w', encoding="utf-8", newline='', errors="replace") as out_file: out_file.write('\n'.join(out_lines))
内容的提问来源于stack exchange,提问作者pglove
相关产品推荐
相关产品推荐

