如何通过Author ID和时间戳用Tweepy/Twitter API V2检索推文?
核心思路
Twitter API v2的/2/users/{id}/tweets端点支持按时间范围过滤推文,结合Tweepy的get_users_tweets方法,你可以针对每个author_id拉取其目标时间戳附近的推文,再通过精确时间匹配找到对应的tweet_id。
关键步骤与代码实现
时间戳格式转换
Twitter API要求时间参数为ISO 8601格式(例如2023-10-05T14:48:00Z),需先将数据集内的时间戳转换为该格式。调用用户推文时间线接口
使用Tweepy的get_users_tweets方法,给目标精确时间前后加1秒缓冲(避免时间精度差异导致漏查),同时请求tweet.created_at字段用于后续匹配。精确匹配时间戳
在返回的推文列表中,找到created_at与目标时间完全一致的条目,提取其id。
示例代码:
import tweepy from datetime import datetime, timedelta # 初始化客户端 client = tweepy.Client(bearer_token, wait_on_rate_limit=True) def get_tweet_id_by_author_and_time(author_id, target_time): # 转换目标时间为ISO格式,并设置1秒缓冲的时间范围 target_dt = datetime.fromisoformat(target_time.replace('Z', '+00:00')) start_time = (target_dt - timedelta(seconds=1)).isoformat().replace('+00:00', 'Z') end_time = (target_dt + timedelta(seconds=1)).isoformat().replace('+00:00', 'Z') # 拉取时间范围内的用户推文,包含创建时间字段 tweets = client.get_users_tweets( id=author_id, start_time=start_time, end_time=end_time, tweet_fields=['created_at'], max_results=100 # 覆盖同秒多发的场景 ) # 遍历匹配精确时间 if tweets.data: for tweet in tweets.data: # 匹配到秒级时间(若数据集精度更高,可调整匹配逻辑) if tweet.created_at.replace(tzinfo=None) == target_dt.replace(tzinfo=None): return tweet.id return None # 示例调用(替换为你数据集里的实际值) author_id = 1241497988423454720 target_time = "2023-10-05T14:48:00Z" tweet_id = get_tweet_id_by_author_and_time(author_id, target_time) print(f"匹配到的推文ID: {tweet_id}")
注意事项
- 时间精度对齐:确保数据集时间戳与Twitter返回的
created_at精度一致(通常为秒级),若存在毫秒差异,需调整匹配逻辑(如仅比较到秒)。 - 同秒多发处理:若同一作者在同一秒发布多条推文,需额外条件(如推文内容片段)辅助定位,可在
get_users_tweets中请求text字段补充匹配。 - 速率限制:Twitter API对该端点有调用频次限制,Tweepy的
wait_on_rate_limit=True会自动等待,但处理百万级数据时需注意批量效率。
内容的提问来源于stack exchange,提问作者Maxl Gemeinderat
相关产品推荐
相关产品推荐

