如何使用tweepy.Paginator从推文中获取Twitter用户及地点数据
问题根因
拿不到关联用户、地点数据的核心错误是使用了Paginator.flatten()遍历结果:
- 这个方法会把接口返回结果做扁平化处理,仅保留主资源(也就是推文对象),直接丢弃
includes字段下所有关联扩展数据,包括你指定的用户、地点信息,自然没法通过tweet.includes访问对应内容。 - 要拿到关联数据,必须直接遍历分页响应对象,逐页同时读取推文数据和includes下的关联资源,自行做ID匹配。
修正步骤
- 移除遍历逻辑里的
flatten()调用,直接迭代Paginator对象拿到每一页的Response实例 - 对每一页响应,分别从
response.data取当前页的推文列表,从response.includes取当前页关联的用户、地点数据 - 构建ID到资源的映射字典:用户ID对应用户信息、地点ID对应地点信息,遍历单条推文时,通过推文自带的
author_id、geo.place_id查映射表就能拿到关联数据 - 原请求的
tweet_fields漏加了geo字段,会导致推文对象上读不到place_id,没法匹配地点数据,需要补上。
修正后可运行代码
import tweepy import time client = tweepy.Client("MYSECRETTOKEN") query = '"Matariki" lang:en' start_time = "2022-06-23T22:00:00.000Z" end_time = "2022-06-26T22:00:00.000Z" paginator = tweepy.Paginator( client.search_all_tweets, query=query, start_time=start_time, end_time=end_time, tweet_fields=['id', 'author_id', "created_at", "text", "source", 'lang', 'in_reply_to_user_id', 'conversation_id', 'public_metrics', 'referenced_tweets', 'reply_settings', 'geo'], user_fields=["name", "username", "location", "verified", "description", "created_at"], place_fields=['full_name', 'id', 'country', 'country_code', 'geo', 'name', 'place_type'], expansions=['author_id', 'geo.place_id'], max_results=10 ) count = 0 for page in paginator: # 构建当前页的ID-资源映射 user_map = {user.id: user for user in page.includes.get('users', [])} place_map = {place.id: place for place in page.includes.get('places', [])} for tweet in page.data: count += 1 # 获取关联作者信息 author = user_map.get(tweet.author_id) # 获取关联地点信息,无地理标签的推文直接返回None place = None if tweet.geo and 'place_id' in tweet.geo: place = place_map.get(tweet.geo['place_id']) # 以下可替换为自定义的数据处理逻辑 print(f"推文{count}内容:{tweet.text}") if author: print(f"发布者:@{author.username} | 认证状态:{author.verified} | 简介:{author.description}") if place: print(f"发布地点:{place.full_name} | 国家编码:{place.country_code}") print('-'*50) # 达到20条上限直接终止 if count >= 20: break # 每10条请求休眠1秒,避免触发接口限流 if count % 10 == 0: time.sleep(1) if count >=20: break
注意事项
- 大部分普通推文不会携带地理标签,因此匹配不到地点数据是正常情况,不属于代码错误
expansions参数传入列表和传入逗号拼接的字符串效果完全一致,列表形式可读性更高- 不要尝试在单条tweet对象上找
includes属性,includes是页级别的响应字段,不属于单条推文的属性。
内容的提问来源于stack exchange,提问作者Martien Lubberink
相关产品推荐
相关产品推荐

