You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Tweepy爬取Twitter推文pandas读CSV报utf-8解码错误如何解决

问题描述

基于Python tweepy库开发的Twitter推文采集爬虫长期用于科研数据支撑,近期运行出现字符解码类故障:

  • 原爬虫写入逻辑:以追加模式打开tweets.csv,遍历采集到的推文时先去除全文首尾空白,转ASCII编码时忽略emoji等Unicode字符,按空白、换行切分文本后过滤@提及与URL链接,拼接回字符串后反转义HTML实体,最终将推文发布时间、属地、处理后文本写入CSV。
  • 原爬虫实现代码:
import csv
import html
# 其余tweepy相关导入省略

# open and create a file to append the data to
csvFile = open('tweets.csv', 'a')
csvWriter = csv.writer(csvFile)
    # use the csv file
    # loop through the tweets variable and add contents to the CSV file
for tweet in tweets:
    text = tweet.full_text.strip()
    #convert the text to ascii ignoring all unicode characters, eg. emojis
    text_ascii = text.encode('ascii','ignore').decode()
    #split the text on whitespace and newlines into a list of words
    text_list = text_ascii.split()
    #iterate over the words, removing @ mentions or URLs 
    text_list_filtered = [word for word in text_list if not (word.startswith('@') or word.startswith('http'))]
    #join the list back into a string
    text_filtered = ' '.join(text_list_filtered)
    #decoding html escaped characters
    text_filtered = html.unescape(text_filtered)
    #write text to the CSV file
    csvWriter.writerow([tweet.created_at, tweet.place, text_filtered])
    print(tweet.created_at, tweet.place, text_filtered)
csvFile.close() 
  • 故障现象:使用pandas读取生成的CSV文件时触发解码报错,报错信息:
UnicodeDecodeError: 'utf-8' codec can't decode byte 0xe1 in position 139390: invalid continuation byte
  • 触发报错的代码行:
tweetsdf = pd.read_csv('tweets.csv')
  • 已尝试无效方案:将编码处理逻辑从text_ascii = text.encode('ascii','ignore').decode()修改为text_ascii = text.encode('utf-8','ignore').decode(),采集时仍出现相同问题。
故障根因
  1. 文件打开未指定编码:Python的open()函数在不同操作系统下默认编码不一致,Windows系统默认使用GBK/CP1252等本地编码写入文件,pandas默认按UTF-8编码读取文件,编码不匹配直接触发解码错误。
  2. 字符处理顺序错误:原代码先做ASCII编码过滤,再执行html.unescape()反转义HTML实体,反转义过程中会生成新的非ASCII字符(比如á转成á,对应报错里的0xe1字节),这些字符没有被过滤就直接写入文件,进一步加剧编码混乱。
  3. 修改后的编码逻辑完全无效:Python3中字符串默认是Unicode类型,对Unicode字符串执行encode('utf-8','ignore').decode()相当于把字符串转成UTF-8字节再原封不动转回来,没有任何过滤非ASCII字符的作用,非ASCII字符会全部保留,再被系统默认编码写入文件,自然还是会报错。
解决方案

1. 修复爬虫写入逻辑,从根源避免编码问题

调整后代码如下,核心修改点:

  • 打开文件时强制指定encoding='utf-8',同时加newline=''避免CSV写入多余空行
  • 调整字符处理顺序:先做HTML实体反转义,再做非ASCII字符过滤
  • 如果不需要保留西语重音、特殊符号等非ASCII内容,保留ASCII转码逻辑即可;如果需要保留这类科研有用的文本内容,直接删掉ASCII转码逻辑,用UTF-8写入不会丢失数据
import csv
import html
# 其余tweepy相关导入省略

# 打开文件时强制指定utf-8编码,加newline参数适配csv模块规范
csvFile = open('tweets_fixed.csv', 'a', encoding='utf-8', newline='')
csvWriter = csv.writer(csvFile)

for tweet in tweets:
    text = tweet.full_text.strip()
    # 先做HTML实体反转义,再做后续处理
    text = html.unescape(text)
    # 转ASCII过滤非ASCII字符(如果需要保留重音等字符,直接删掉下面这行即可)
    text_ascii = text.encode('ascii', 'ignore').decode()
    # 切分、过滤@和URL
    text_list = text_ascii.split()
    text_list_filtered = [word for word in text_list if not (word.startswith('@') or word.startswith('http'))]
    text_filtered = ' '.join(text_list_filtered)
    # 写入文件
    csvWriter.writerow([tweet.created_at, tweet.place, text_filtered])
    print(tweet.created_at, tweet.place, text_filtered)

csvFile.close()

2. 读取已损坏的历史CSV文件

对于之前已经生成的、编码混乱的tweets.csv,可以用以下两种方式读取:

  • 方式一:按Latin-1编码读取,该编码支持单字节全匹配,不会出现解码报错,适合混了CP1252/UTF-8内容的旧文件:
import pandas as pd
tweetsdf = pd.read_csv('tweets.csv', encoding='latin-1')
  • 方式二:仍按UTF-8读取,自动替换/忽略无法解码的坏字节:
import pandas as pd
# 坏字节替换为�占位符
tweetsdf = pd.read_csv('tweets.csv', encoding='utf-8', encoding_errors='replace')
# 或者直接跳过坏字节
# tweetsdf = pd.read_csv('tweets.csv', encoding='utf-8', encoding_errors='ignore')

内容的提问来源于stack exchange,提问作者Shehzadi Aziz

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.27 19:01:47