如何解决Python 3写入文件时商标符号引发的cp950编码错误?
解决Python写入含特殊字符文本时的Unicode编码错误
你遇到的问题核心在于默认文件编码不支持Unicode特殊字符——你的系统默认用了cp950编码(繁体中文常用编码),它没法处理像™(\u2122)这类Unicode字符,所以写入时直接报错了。下面给你几个靠谱的解决办法:
方法1:打开文件时指定UTF-8编码(最推荐)
这是最简单直接的方案,打开文件时明确指定编码为utf-8,就能完美支持所有Unicode字符:
import requests from inscriptis import get_text from bs4 import BeautifulSoup page = requests.get(r'http://www3.asiainsurancereview.com//News/View-NewsLetter-Article/id/42528/Type/eDaily/Technology-First-round-of-the-pre-launch-of-the-Ydentity-ICO-starts-today') soup = BeautifulSoup(page.text, 'lxml') html = soup.find(class_='article-wrap') text = get_text(html.text) print(text) # 关键:打开文件时指定encoding='utf-8' with open('test.txt', 'w', encoding='utf-8') as articleFile: articleFile.write(text)
这里用with语句还能自动关闭文件,比手动close()更安全省心。
方法2:正确处理编码转换(如果需要字节写入)
如果你之前尝试了encode("utf-8"),那打开文件时要改成二进制写入模式wb,因为encode后得到的是bytes类型,只能写入二进制模式的文件:
text = get_text(html.text) text_bytes = text.encode("utf-8") with open('test.txt', 'wb') as articleFile: articleFile.write(text_bytes)
不过这种方法不如方法1直观,一般优先选方法1。
方法3:用unidecode转换特殊字符为ASCII兼容形式
你之前用unidecode的方式出错,是因为text已经是str类型了,不需要再用str(text, encoding="utf-8")解码(字符串本身不能再解码)。正确的用法是直接把str传给unidecode:
import requests from inscriptis import get_text from bs4 import BeautifulSoup from unidecode import unidecode page = requests.get(r'http://www3.asiainsurancereview.com//News/View-NewsLetter-Article/id/42528/Type/eDaily/Technology-First-round-of-the-pre-launch-of-the-Ydentity-ICO-starts-today') soup = BeautifulSoup(page.text, 'lxml') html = soup.find(class_='article-wrap') text = get_text(html.text) # 直接传入str即可 processed_text = unidecode(text) with open('test.txt', 'w') as articleFile: articleFile.write(processed_text)
这样会把™转换成(TM)这类ASCII兼容的字符,就不会触发编码错误了。
为什么之前的判断字符串类型方法没用?
你最后那个判断类型的方法没解决问题,是因为问题出在写入时的文件编码,而不是字符串本身的类型——text已经是str了,但写入时用默认的cp950编码还是没法处理特殊字符,所以指定文件编码才是关键。
内容的提问来源于stack exchange,提问作者Kristada673
相关产品推荐
相关产品推荐

