Tweepy爬取Twitter推文pandas读CSV报utf-8解码错误如何解决
问题描述
基于Python tweepy库开发的Twitter推文采集爬虫长期用于科研数据支撑,近期运行出现字符解码类故障:
- 原爬虫写入逻辑:以追加模式打开
tweets.csv,遍历采集到的推文时先去除全文首尾空白,转ASCII编码时忽略emoji等Unicode字符,按空白、换行切分文本后过滤@提及与URL链接,拼接回字符串后反转义HTML实体,最终将推文发布时间、属地、处理后文本写入CSV。 - 原爬虫实现代码:
import csv import html # 其余tweepy相关导入省略 # open and create a file to append the data to csvFile = open('tweets.csv', 'a') csvWriter = csv.writer(csvFile) # use the csv file # loop through the tweets variable and add contents to the CSV file for tweet in tweets: text = tweet.full_text.strip() #convert the text to ascii ignoring all unicode characters, eg. emojis text_ascii = text.encode('ascii','ignore').decode() #split the text on whitespace and newlines into a list of words text_list = text_ascii.split() #iterate over the words, removing @ mentions or URLs text_list_filtered = [word for word in text_list if not (word.startswith('@') or word.startswith('http'))] #join the list back into a string text_filtered = ' '.join(text_list_filtered) #decoding html escaped characters text_filtered = html.unescape(text_filtered) #write text to the CSV file csvWriter.writerow([tweet.created_at, tweet.place, text_filtered]) print(tweet.created_at, tweet.place, text_filtered) csvFile.close()
- 故障现象:使用pandas读取生成的CSV文件时触发解码报错,报错信息:
UnicodeDecodeError: 'utf-8' codec can't decode byte 0xe1 in position 139390: invalid continuation byte
- 触发报错的代码行:
tweetsdf = pd.read_csv('tweets.csv')
- 已尝试无效方案:将编码处理逻辑从
text_ascii = text.encode('ascii','ignore').decode()修改为text_ascii = text.encode('utf-8','ignore').decode(),采集时仍出现相同问题。
故障根因
- 文件打开未指定编码:Python的
open()函数在不同操作系统下默认编码不一致,Windows系统默认使用GBK/CP1252等本地编码写入文件,pandas默认按UTF-8编码读取文件,编码不匹配直接触发解码错误。 - 字符处理顺序错误:原代码先做ASCII编码过滤,再执行
html.unescape()反转义HTML实体,反转义过程中会生成新的非ASCII字符(比如á转成á,对应报错里的0xe1字节),这些字符没有被过滤就直接写入文件,进一步加剧编码混乱。 - 修改后的编码逻辑完全无效:Python3中字符串默认是Unicode类型,对Unicode字符串执行
encode('utf-8','ignore').decode()相当于把字符串转成UTF-8字节再原封不动转回来,没有任何过滤非ASCII字符的作用,非ASCII字符会全部保留,再被系统默认编码写入文件,自然还是会报错。
解决方案
1. 修复爬虫写入逻辑,从根源避免编码问题
调整后代码如下,核心修改点:
- 打开文件时强制指定
encoding='utf-8',同时加newline=''避免CSV写入多余空行 - 调整字符处理顺序:先做HTML实体反转义,再做非ASCII字符过滤
- 如果不需要保留西语重音、特殊符号等非ASCII内容,保留ASCII转码逻辑即可;如果需要保留这类科研有用的文本内容,直接删掉ASCII转码逻辑,用UTF-8写入不会丢失数据
import csv import html # 其余tweepy相关导入省略 # 打开文件时强制指定utf-8编码,加newline参数适配csv模块规范 csvFile = open('tweets_fixed.csv', 'a', encoding='utf-8', newline='') csvWriter = csv.writer(csvFile) for tweet in tweets: text = tweet.full_text.strip() # 先做HTML实体反转义,再做后续处理 text = html.unescape(text) # 转ASCII过滤非ASCII字符(如果需要保留重音等字符,直接删掉下面这行即可) text_ascii = text.encode('ascii', 'ignore').decode() # 切分、过滤@和URL text_list = text_ascii.split() text_list_filtered = [word for word in text_list if not (word.startswith('@') or word.startswith('http'))] text_filtered = ' '.join(text_list_filtered) # 写入文件 csvWriter.writerow([tweet.created_at, tweet.place, text_filtered]) print(tweet.created_at, tweet.place, text_filtered) csvFile.close()
2. 读取已损坏的历史CSV文件
对于之前已经生成的、编码混乱的tweets.csv,可以用以下两种方式读取:
- 方式一:按Latin-1编码读取,该编码支持单字节全匹配,不会出现解码报错,适合混了CP1252/UTF-8内容的旧文件:
import pandas as pd tweetsdf = pd.read_csv('tweets.csv', encoding='latin-1')
- 方式二:仍按UTF-8读取,自动替换/忽略无法解码的坏字节:
import pandas as pd # 坏字节替换为�占位符 tweetsdf = pd.read_csv('tweets.csv', encoding='utf-8', encoding_errors='replace') # 或者直接跳过坏字节 # tweetsdf = pd.read_csv('tweets.csv', encoding='utf-8', encoding_errors='ignore')
内容的提问来源于stack exchange,提问作者Shehzadi Aziz
相关产品推荐
相关产品推荐

