如何用Tweepy(学术权限)获取指定时段的随机10万条Twitter推文?
问题
我目前使用带有学术访问权限的Tweepy获取指定时间区间内的Twitter推文。采用通用查询,需要10万条该时段内的随机推文,但当前代码返回的是该时段内最新的10万条。请问能否实现需求?若可行,应如何修改代码?
当前获取最新推文的代码:
# Imports import tweepy import json import csv # Store bearer_token in variable bearer_token = "Input Bearer Token Here" client = tweepy.Client(bearer_token=bearer_token) # Replace with your own search query query = ' "Air Pollution" "2.5" place_country:US' # Replace with time period of your choice start_time = '2021-11-20T00:00:00Z' # Replace with time period of your choice end_time = '2022-11-20T00:00:00Z' tweets = client.search_all_tweets(query=query, tweet_fields=['context_annotations', 'created_at', 'geo'], place_fields = ['place_type','geo'], expansions='geo.place_id', start_time=start_time, end_time=end_time, max_results=100000) # Prepare to write to csv file f = open('tweetData.csv','w') writer = csv.writer(f) # Write to csv file for tweet in tweets.data: print(tweet.text) print(tweet.created_at) writer.writerow(['0', tweet.id, tweet.created_at, tweet.text]) # Close csv file f.close()
解决方案
Twitter API v2的search_all_tweets接口本身没有直接返回随机推文的参数,默认按时间倒序(最新优先)返回结果。要实现随机采样,有两种可行方案:
方案一:分时间段随机采样
把目标时间区间拆分成多个子时间段,随机选择子时间段查询,直到凑够10万条推文。这种方法能均匀覆盖整个时间区间,避免只取最新内容。
修改步骤
- 导入
random和datetime模块,用于生成随机子时间段 - 将原时间区间拆分为更小的单位(比如按天拆分)
- 打乱子时间段顺序,逐个调用接口获取推文,直到收集满10万条
修改后代码
# Imports import tweepy import csv import random from datetime import datetime, timedelta # Store bearer_token in variable bearer_token = "Input Bearer Token Here" client = tweepy.Client(bearer_token=bearer_token) # Replace with your own search query query = ' "Air Pollution" "2.5" place_country:US' # Parse original time period start_datetime = datetime.fromisoformat(start_time.replace('Z', '+00:00')) end_datetime = datetime.fromisoformat(end_time.replace('Z', '+00:00')) # Split into daily intervals time_intervals = [] current_start = start_datetime while current_start < end_datetime: current_end = current_start + timedelta(days=1) if current_end > end_datetime: current_end = end_datetime time_intervals.append(( current_start.isoformat().replace('+00:00', 'Z'), current_end.isoformat().replace('+00:00', 'Z') )) current_start = current_end # Randomize the order of intervals random.shuffle(time_intervals) # Collect tweets until we reach 100k collected_tweets = [] max_needed = 100000 for interval_start, interval_end in time_intervals: if len(collected_tweets) >= max_needed: break # Use paginator to fetch all tweets in this interval (max 100 per call) paginator = tweepy.Paginator( client.search_all_tweets, query=query, tweet_fields=['context_annotations', 'created_at', 'geo'], place_fields=['place_type','geo'], expansions='geo.place_id', start_time=interval_start, end_time=interval_end, max_results=100 ) for response in paginator: if response.data: collected_tweets.extend(response.data) if len(collected_tweets) >= max_needed: collected_tweets = collected_tweets[:max_needed] break # Write to CSV with open('tweetData.csv','w', newline='') as f: writer = csv.writer(f) writer.writerow(['index', 'tweet_id', 'created_at', 'text']) for idx, tweet in enumerate(collected_tweets): writer.writerow([idx+1, tweet.id, tweet.created_at, tweet.text])
方案二:批量获取后随机抽取
如果目标时间区间内的推文总数远大于10万,可以先批量获取足够多的推文(比如15万条),再从中随机抽取10万条。若推文总数不足10万,则直接取全部。
修改步骤
- 使用分页获取远超10万条的推文
- 用
random.sample从结果中随机抽取目标数量
修改后代码片段
# ... 保留原导入和client初始化代码 ... # Fetch enough tweets with pagination all_tweets = [] paginator = tweepy.Paginator( client.search_all_tweets, query=query, tweet_fields=['context_annotations', 'created_at', 'geo'], place_fields=['place_type','geo'], expansions='geo.place_id', start_time=start_time, end_time=end_time, max_results=100 ) for response in paginator: if response.data: all_tweets.extend(response.data) # Stop when we have enough buffer if len(all_tweets) >= 150000: break # Randomly sample 100k tweets sampled_tweets = random.sample(all_tweets, 100000) if len(all_tweets) >= 100000 else all_tweets # Write to CSV with open('tweetData.csv','w', newline='') as f: writer = csv.writer(f) writer.writerow(['index', 'tweet_id', 'created_at', 'text']) for idx, tweet in enumerate(sampled_tweets): writer.writerow([idx+1, tweet.id, tweet.created_at, tweet.text])
注意事项
- 学术访问权限有API调用额度限制,批量查询时需注意不要超额
- 方案一的随机性更均匀,适合覆盖全时间段;方案二更简洁,但依赖于目标时段有足够多的推文
- 原代码中
max_results=100000无效,接口单次调用最多返回100条,必须用tweepy.Paginator实现批量获取
内容的提问来源于stack exchange,提问作者Rashid Abramov
相关产品推荐
相关产品推荐

