如何通过Python Tweepy高效批量获取我关注的所有用户的推文?
Great question—this is a super common pain point I’ve dealt with a lot when building tools on Twitter’s API. Making individual calls for every followee eats up rate limits and takes forever, so let’s break down the most efficient ways to handle this with Tweepy, focusing on the modern API v2 (since v1.1 is on its way out):
1. 先批量拉取关注用户ID(砍初始调用量)
First, don’t waste calls fetching followees one by one. Use Tweepy’s bulk endpoint to grab all your followee IDs in a single paginated request—this cuts your initial call count from hundreds to just a handful.
示例代码:
import tweepy # 初始化Tweepy Client(替换成你的API密钥) client = tweepy.Client( bearer_token="YOUR_BEARER_TOKEN", consumer_key="YOUR_CONSUMER_KEY", consumer_secret="YOUR_CONSUMER_SECRET", access_token="YOUR_ACCESS_TOKEN", access_token_secret="YOUR_ACCESS_TOKEN_SECRET", wait_on_rate_limit=True # 自动处理速率限制,省得自己写重试逻辑 ) # 分页获取当前账号的全部关注用户ID followee_ids = [] current_user_id = client.get_me().data.id for page in tweepy.Paginator( client.get_users_following, id=current_user_id, max_results=1000 # 每次请求最多返回1000个ID,拉满效率 ): if page.data: followee_ids.extend([user.id for user in page.data]) print(f"搞定!共获取到 {len(followee_ids)} 个关注用户ID")
2. 批量获取推文:两种核心方案
方案A:用搜索端点一次性过滤多个用户
Twitter API v2的搜索端点允许你用from:user_id语法批量查询多个用户的推文,把N个用户的查询合并成一个(或少量)调用。注意查询字符串有512字符的限制,所以要把followee IDs分成合适的批次(比如每批次30-40个ID,取决于ID长度)。
示例代码:
def batch_fetch_recent_tweets(followee_ids, batch_size=35): # 分批次处理关注用户,避免查询字符串过长 for i in range(0, len(followee_ids), batch_size): batch = followee_ids[i:i+batch_size] # 构建批量查询语句:from:id1 OR from:id2 OR ... query = " OR ".join([f"from:{user_id}" for user_id in batch]) # 分页拉取该批次用户的所有近期推文 for tweet_page in tweepy.Paginator( client.search_recent_tweets, query=query, tweet_fields=["created_at", "author_id"], max_results=100 ): if tweet_page.data: for tweet in tweet_page.data: # 这里替换成你自己的处理逻辑,比如存数据库或打印 print(f"@{tweet.author_id}: {tweet.text[:50]}... ({tweet.created_at})") # 启动批量拉取 batch_fetch_recent_tweets(followee_ids)
优点:大幅减少API调用次数;返回的推文按时间统一排序,方便批量处理。
缺点:只能获取最近7天的推文(除非你有学术研究访问权限);无法获取受保护用户的推文。
方案B:并行处理单个用户请求(适合拉取全量历史推文)
如果你需要获取用户的全部历史推文(而不仅仅是近期的),搜索端点就不够用了。这时可以用Python的多线程并行调用get_users_tweets,同时严格遵守速率限制,比串行快好几倍。
示例代码:
from concurrent.futures import ThreadPoolExecutor def fetch_user_full_tweets(user_id): # 分页拉取单个用户的全部推文 user_tweets = [] for page in tweepy.Paginator( client.get_users_tweets, id=user_id, tweet_fields=["created_at"], max_results=100 ): if page.data: user_tweets.extend(page.data) return user_id, user_tweets # 用线程池并行处理(线程数别设太高,建议5-10,避免触发API限制) with ThreadPoolExecutor(max_workers=8) as executor: # 提交所有用户的推文拉取任务 futures = [executor.submit(fetch_user_full_tweets, uid) for uid in followee_ids] # 处理返回结果 for future in futures: user_id, tweets = future.result() print(f"用户 {user_id} 共有 {len(tweets)} 条历史推文") # 这里添加存储或处理推文的逻辑
优点:能获取用户的全量历史推文(只要API允许);可以处理受保护用户的推文(如果用户授权)。
缺点:并行请求需要控制速率,避免触发API限制;总调用次数仍和用户数量相关,但比串行高效很多。
关键优化技巧
- 一定要开
wait_on_rate_limit=True:Trust me,这能省掉你写重试逻辑的大把时间,Tweepy会自动检测速率限制并等待重置。 - 只请求需要的字段:在
tweet_fields里只指定你需要的字段(比如created_at、author_id),减少数据传输量,提升速度。 - 缓存已处理的用户数据:如果需要重复运行脚本,缓存已经拉取过的用户推文,避免重复调用API。
- 优先用API v2:v2的速率限制更宽松,批量支持也更友好,v1.1已经逐步被淘汰了。
内容的提问来源于stack exchange,提问作者user12377624

