为何仅使用Python的print语句时会触发UnicodeEncodeError?
为什么Python中print()函数会抛出UnicodeEncodeError?
我近期开始学习Python,为练习开发了「常用词查找工具」。该工具爬取Jisho网站的#kanji页面,提取ul类为no-bullet的音读、训读词组,目标是找出并打印最常见英文单词。使用VS Code作为IDE,已导入urllib.parse、requests和bs4的BeautifulSoup,代码如下:
kanji = '人' parsed_kanji = urllib.parse.quote(kanji) url = f'https://jisho.org/search/{parsed_kanji}%20%23kanji' page = requests.get(url) soup = BeautifulSoup(page.text, "html.parser") compounds = [] for li in soup.select('.no-bullet li'): comp = ' '.join(li.text.split()) compounds.append(comp) print(compounds)
(查找最常见单词的代码未包含)
当代码无print(compounds)时运行正常,但添加该语句后出现如下错误:
Traceback (most recent call last): File "c:\Users\Lugnut\OneDrive\Desktop\frequent\most_common\test_list.py", line 22, in <module> print(compounds) File "C:\Program Files\WindowsApps\PythonSoftwareFoundation.Python.3.9_3.9.3568.0_x64__qbz5n2kfra8p0\lib\encodings\cp1252.py", line 19, in encode return codecs.charmap_encode(input,self.errors,encoding_table)[0] UnicodeEncodeError: 'charmap' codec can't encode character '\u4eba' in position 2: character maps to <undefined>
问题原因
这是Windows终端默认采用的CP1252编码不支持Unicode字符(比如你爬取到的汉字「人」,对应Unicode编码\u4eba)导致的。print()函数输出时会尝试用终端的默认编码对内容进行编码,但CP1252编码表中没有对应汉字的映射关系,因此抛出编码错误。
解决方法
方法1:修改标准输出编码
在代码开头添加以下代码,强制将标准输出的编码设置为UTF-8:import sys sys.stdout.reconfigure(encoding='utf-8')方法2:print时指定错误处理策略
在print()中通过errors参数指定无法编码时的处理方式:print(compounds, errors='ignore') # 忽略无法编码的字符 # 或 print(compounds, errors='replace') # 用?替换无法编码的字符方法3:切换终端编码为UTF-8
在VS Code中,可将终端默认配置切换为Git Bash或PowerShell,这类终端默认支持UTF-8编码,能正常显示Unicode字符。
内容的提问来源于stack exchange,提问作者Lugnut
相关产品推荐
相关产品推荐

