You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何复用Twitter API搜索结果?排查读取文件UnicodeDecodeError问题

Twitter数据读取的UnicodeDecodeError及正确保存方法

错误原因

  1. 编码不匹配:手动复制输出到记事本保存时,Windows记事本默认使用cp1252编码,但Twitter返回的印尼语数据包含cp1252不支持的特殊字符,读取时Python默认用系统编码(即cp1252)解码,触发UnicodeDecodeError。
  2. 数据格式问题:手动复制的是tweepy.models.SearchResults对象的打印字符串,这是对象的非结构化表示,不仅包含推文内容,还有大量对象元数据,后续处理会非常麻烦。

正确保存Twitter搜索结果的方法

方法1:保存为JSON格式(推荐,结构化易处理)

tweepy的Status对象自带_json属性,可直接获取原始JSON结构数据,保存为UTF-8编码的JSON文件:

import tweepy
import json

# 认证代码
api_key = "CENSORED"
api_key_secret = "CENSORED"
access_token = "CENSORED"
access_token_secret = "CENSORED"

auth = tweepy.OAuthHandler(api_key, api_key_secret)
auth.set_access_token(access_token, access_token_secret)
api = tweepy.API(auth)

# 获取搜索结果
hasilSearch = api.search_tweets(
    q="(rokok OR merokok OR nyebat OR ngerokok OR ngudut OR ngudud OR udud OR sebat OR sebats OR roko OR cerutu OR cangklong OR tembakau) (kanker OR tumor OR ganas)",
    lang='id',
    count=10000
)

# 提取每条推文的JSON数据
tweets_list = [tweet._json for tweet in hasilSearch]

# 保存为UTF-8编码的JSON文件
with open('20230227_tweets.json', 'w', encoding='utf-8') as f:
    json.dump(tweets_list, f, ensure_ascii=False, indent=2)

读取JSON文件的代码:

import json

with open('20230227_tweets.json', 'r', encoding='utf-8') as f:
    tweets_data = json.load(f)

# 示例:遍历并打印每条推文文本
for tweet in tweets_data:
    print(tweet['text'])

方法2:保存为纯文本文件(仅需推文内容时)

如果只需要保存推文文本,直接用代码写入UTF-8编码的文本文件,避免手动复制的编码问题:

# 获取hasilSearch之后执行
with open('20230227.txt', 'w', encoding='utf-8') as f:
    for tweet in hasilSearch:
        f.write(tweet.text + '\n')

读取文本文件的代码:

with open('20230227.txt', 'r', encoding='utf-8') as f:
    content = f.read()
print(content)

内容的提问来源于stack exchange,提问作者Ibrahim Dharmawan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.30 07:18:00