Twitter API v2与Tweepy如何获取流式推文的作者信息
问题根源
Twitter API v2 流式接口默认仅返回推文text等极少数字段,created_at、author_id、作者信息等都属于需要主动申请的扩展字段,不在默认返回范围内;同时v2版本将推文、用户等实体拆分为独立对象,不会直接把作者信息挂在推文对象下,直接读取就会得到None。
修复方法
- 启动流时,通过
filter()方法的参数明确声明你需要的推文字段、用户字段,以及要关联拉取的作者扩展 - 重写
on_response()方法替代仅拿推文对象的on_tweet(),从完整响应里同时拿到推文数据和关联的作者数据
可直接运行的修正代码
import tweepy import time import pandas as pd # 替换为你自己的Twitter API v2 Bearer Token BEARER_TOKEN = "你的Bearer Token" class MyStream(tweepy.StreamingClient): def on_connect(self): print("Connected") def on_response(self, response): tweet = response.data # 构建用户ID到用户对象的映射 user_map = {user.id: user for user in response.includes.get("users", [])} # 过滤转发、引用、回复推文,和原有逻辑一致 if tweet.referenced_tweets is None: # 之前拿不到的字段现在可以正常读取 print(f"推文ID: {tweet.id}") print(f"发布时间: {tweet.created_at}") print(f"作者ID: {tweet.author_id}") # 从映射里取作者信息 author = user_map.get(tweet.author_id) if author: # author.username就是你要的user_screenname print(f"作者用户名: {author.username}") print(f"作者显示名: {author.name}") print(f"推文内容: {tweet.text}") print("-" * 60) # 存CSV的逻辑可以直接用 # data.append([tweet.created_at, author.username, tweet.text]) # pd.DataFrame(data, columns=columns).to_csv("StreamTweets.csv", index=False) time.sleep(0.1) if __name__ == "__main__": stream = MyStream(bearer_token=BEARER_TOKEN) # 清理已有流规则,避免规则重复生效 old_rules = stream.get_rules().data if old_rules: stream.delete_rules([rule.id for rule in old_rules]) # 添加你需要的过滤规则,示例为过滤带#Python话题标签的推文 stream.add_rules(tweepy.StreamRule("#Python")) # 启动流时必须传入你要的字段和扩展配置 stream.filter( # 声明需要的推文额外字段 tweet_fields=["created_at", "author_id", "id"], # 声明需要的用户字段 user_fields=["username", "name"], # 声明要同步返回推文关联的作者用户对象 expansions=["author_id"] )
注意事项
- 不要尝试在
on_tweet()回调里直接读取tweet.user、tweet.user_screenname这类属性,v2接口不会在推文对象下嵌套用户实体,必须从响应的includes字段里取关联的用户数据 - 所有非默认返回的字段,都必须在
filter()的对应参数里明确列出,否则接口不会返回对应数据,读取结果就是None - 点赞、发推等写操作不能使用Bearer Token身份调用,需要换成带用户上下文授权的Client实例
内容的提问来源于stack exchange,提问作者Nils Gallist
相关产品推荐
相关产品推荐

