You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Tweepy(学术权限)获取指定时段的随机10万条Twitter推文?

问题

我目前使用带有学术访问权限的Tweepy获取指定时间区间内的Twitter推文。采用通用查询,需要10万条该时段内的随机推文,但当前代码返回的是该时段内最新的10万条。请问能否实现需求?若可行,应如何修改代码?

当前获取最新推文的代码:

# Imports
import tweepy
import json
import csv

# Store bearer_token in variable
bearer_token = "Input Bearer Token Here"

client = tweepy.Client(bearer_token=bearer_token)

# Replace with your own search query
query = ' "Air Pollution" "2.5" place_country:US'           

# Replace with time period of your choice
start_time = '2021-11-20T00:00:00Z'

# Replace with time period of your choice
end_time = '2022-11-20T00:00:00Z'

tweets = client.search_all_tweets(query=query, tweet_fields=['context_annotations', 'created_at', 'geo'], 
                                  
                                  place_fields = ['place_type','geo'], expansions='geo.place_id',
                                  start_time=start_time,
                                  end_time=end_time, max_results=100000)

# Prepare to write to csv file
f = open('tweetData.csv','w')
writer = csv.writer(f)

# Write to csv file
for tweet in tweets.data:
    print(tweet.text)
    print(tweet.created_at)
    writer.writerow(['0', tweet.id, tweet.created_at, tweet.text])

# Close csv file
f.close()

解决方案

Twitter API v2的search_all_tweets接口本身没有直接返回随机推文的参数,默认按时间倒序(最新优先)返回结果。要实现随机采样,有两种可行方案:

方案一:分时间段随机采样

把目标时间区间拆分成多个子时间段,随机选择子时间段查询,直到凑够10万条推文。这种方法能均匀覆盖整个时间区间,避免只取最新内容。

修改步骤

  1. 导入random和datetime模块,用于生成随机子时间段
  2. 将原时间区间拆分为更小的单位(比如按天拆分)
  3. 打乱子时间段顺序,逐个调用接口获取推文,直到收集满10万条

修改后代码

# Imports
import tweepy
import csv
import random
from datetime import datetime, timedelta

# Store bearer_token in variable
bearer_token = "Input Bearer Token Here"

client = tweepy.Client(bearer_token=bearer_token)

# Replace with your own search query
query = ' "Air Pollution" "2.5" place_country:US'           

# Parse original time period
start_datetime = datetime.fromisoformat(start_time.replace('Z', '+00:00'))
end_datetime = datetime.fromisoformat(end_time.replace('Z', '+00:00'))

# Split into daily intervals
time_intervals = []
current_start = start_datetime
while current_start < end_datetime:
    current_end = current_start + timedelta(days=1)
    if current_end > end_datetime:
        current_end = end_datetime
    time_intervals.append((
        current_start.isoformat().replace('+00:00', 'Z'), 
        current_end.isoformat().replace('+00:00', 'Z')
    ))
    current_start = current_end

# Randomize the order of intervals
random.shuffle(time_intervals)

# Collect tweets until we reach 100k
collected_tweets = []
max_needed = 100000

for interval_start, interval_end in time_intervals:
    if len(collected_tweets) >= max_needed:
        break
    # Use paginator to fetch all tweets in this interval (max 100 per call)
    paginator = tweepy.Paginator(
        client.search_all_tweets,
        query=query,
        tweet_fields=['context_annotations', 'created_at', 'geo'],
        place_fields=['place_type','geo'],
        expansions='geo.place_id',
        start_time=interval_start,
        end_time=interval_end,
        max_results=100
    )
    for response in paginator:
        if response.data:
            collected_tweets.extend(response.data)
            if len(collected_tweets) >= max_needed:
                collected_tweets = collected_tweets[:max_needed]
                break

# Write to CSV
with open('tweetData.csv','w', newline='') as f:
    writer = csv.writer(f)
    writer.writerow(['index', 'tweet_id', 'created_at', 'text'])
    for idx, tweet in enumerate(collected_tweets):
        writer.writerow([idx+1, tweet.id, tweet.created_at, tweet.text])

方案二:批量获取后随机抽取

如果目标时间区间内的推文总数远大于10万,可以先批量获取足够多的推文(比如15万条),再从中随机抽取10万条。若推文总数不足10万,则直接取全部。

修改步骤

  1. 使用分页获取远超10万条的推文
  2. 用random.sample从结果中随机抽取目标数量

修改后代码片段

# ... 保留原导入和client初始化代码 ...

# Fetch enough tweets with pagination
all_tweets = []
paginator = tweepy.Paginator(
    client.search_all_tweets,
    query=query,
    tweet_fields=['context_annotations', 'created_at', 'geo'],
    place_fields=['place_type','geo'],
    expansions='geo.place_id',
    start_time=start_time,
    end_time=end_time,
    max_results=100
)
for response in paginator:
    if response.data:
        all_tweets.extend(response.data)
        # Stop when we have enough buffer
        if len(all_tweets) >= 150000:
            break

# Randomly sample 100k tweets
sampled_tweets = random.sample(all_tweets, 100000) if len(all_tweets) >= 100000 else all_tweets

# Write to CSV
with open('tweetData.csv','w', newline='') as f:
    writer = csv.writer(f)
    writer.writerow(['index', 'tweet_id', 'created_at', 'text'])
    for idx, tweet in enumerate(sampled_tweets):
        writer.writerow([idx+1, tweet.id, tweet.created_at, tweet.text])

注意事项

  • 学术访问权限有API调用额度限制,批量查询时需注意不要超额
  • 方案一的随机性更均匀,适合覆盖全时间段;方案二更简洁,但依赖于目标时段有足够多的推文
  • 原代码中max_results=100000无效,接口单次调用最多返回100条,必须用tweepy.Paginator实现批量获取

内容的提问来源于stack exchange,提问作者Rashid Abramov

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.11 05:10:28