使用Tweepy StreamingClient保存JSON格式推文至TXT文件遇序列化错误求助
解决Tweepy流推文无法序列化为JSON的问题
你遇到的问题核心是:尽管给StreamingClient设置了return_type=dict,但on_tweet回调接收的依然是Tweet对象而非字典——这个参数主要影响Tweepy客户端的普通API调用返回格式,对流回调的对象类型不起作用。
要解决JSON序列化错误,你需要把Tweet对象转换成可序列化的字典,有两种可靠方式:
方法1:使用Pydantic的model_dump()(Tweepy v4.10及以上版本)
Tweepy的Tweet类基于Pydantic模型,直接调用model_dump()就能得到结构化字典:
import tweepy import json from tweepy import StreamingClient, StreamRule class TweetPrinter(tweepy.StreamingClient): def on_tweet(self, tweet): with open("fetched_tweets.txt", "a") as f: # 将Tweet对象转为字典后序列化,加换行符方便后续读取 f.write(json.dumps(tweet.model_dump(), indent=4) + "\n") return True printer = TweetPrinter(bearer_token=bearer_token) rule = StreamRule(value="Python") printer.add_rules(rule) printer.filter(expansions=['author_id', 'geo.place_id'], tweet_fields="created_at")
方法2:使用_json属性(兼容旧版Tweepy)
如果你的Tweepy版本低于v4.10,可以直接访问Tweet对象的_json属性获取原始响应字典:
def on_tweet(self, tweet): with open("fetched_tweets.txt", "a") as f: f.write(json.dumps(tweet._json, indent=4) + "\n") return True
额外说明:获取完整的扩展数据
如果你需要expansions参数返回的作者、地点等附加数据,仅重写on_tweet不够——因为on_tweet只传递推文主体。此时应该重写on_response方法,获取包含includes字段的完整响应:
class TweetPrinter(tweepy.StreamingClient): def on_response(self, response): # response包含tweet、includes等所有数据,转成字典后序列化 with open("fetched_tweets.txt", "a") as f: f.write(json.dumps(response.model_dump(), indent=4) + "\n") return True
内容的提问来源于stack exchange,提问作者user3262756
相关产品推荐
相关产品推荐

