Python单词频次统计报错UnicodeDecodeError求助(IDLE 3.5.1)
嘿,这个问题我太熟悉啦!你碰到的UnicodeDecodeError是因为你的文本文件根本不是UTF-8编码的,但Python 3里open()函数默认会用UTF-8去解码,碰到不兼容的字节就直接报错了。另外还要注意,你代码里的print k, v在Python 3.5里是语法错误,得改成带括号的形式才行~
咱们一步步来解决:
1. 先解决编码问题
首先得给open()指定正确的文件编码。常见的非UTF-8编码有gbk(中文Windows常用)、cp1252(英文Windows常用)、utf-8-sig(带BOM的UTF-8文件)。你可以挨个试试,比如先试带BOM的UTF-8:
file = open('IntroductoryCS.txt', encoding='utf-8-sig') wordcount = {} for word in file.read().split(): if word not in wordcount: wordcount[word] = 1 else: wordcount[word] += 1 # Python3里print是函数,必须加括号 for k, v in wordcount.items(): print(k, v)
如果还是报错,把encoding='utf-8-sig'换成encoding='gbk'或者encoding='cp1252'再试。
2. 不知道编码?用工具检测
要是实在不知道文件用的啥编码,可以用chardet库来自动检测:
- 先打开命令行(或者IDLE的Shell),输入
pip install chardet安装这个库 - 然后运行下面的代码:
import chardet with open('IntroductoryCS.txt', 'rb') as f: detect_result = chardet.detect(f.read()) print("文件编码是:", detect_result['encoding'])
把输出的编码值放到open()的encoding参数里,就能正常读取啦!
另外,推荐你用with语句打开文件,这样能自动帮你关闭文件,更安全,还能简化统计逻辑:
with open('IntroductoryCS.txt', encoding='你检测到的编码') as file: wordcount = {} for word in file.read().split(): # 用get方法简化if-else判断 wordcount[word] = wordcount.get(word, 0) + 1 for k, v in wordcount.items(): print(k, v)
内容的提问来源于stack exchange,提问作者user9174392
相关产品推荐
相关产品推荐

