You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Tweepy多次调用Twitter API获取单用户超200条推文并保留DataFrame格式

How to Fetch More Than 200 Tweets per Handle with Tweepy & Pandas

Great question! The key here is to use pagination with the Twitter API's max_id parameter to keep fetching older tweets beyond the initial 200. Since you already have wait_on_rate_limit=True set, Tweepy will automatically pause when you hit the rate limit and resume once the window resets—perfect for making those 180 calls in a 15-minute window.

Here's how to modify your code to collect as many tweets as possible (or up to a limit you define) while maintaining your pandas DataFrame structure:

Step 1: Fix the Missing Import

First, notice your original code uses np but doesn't import NumPy—let's add that at the top to avoid errors.

Step 2: Update the get_tweets Function for Pagination

We'll modify the function to loop, fetching batches of 200 tweets each time, using max_id to pull older tweets than the last batch. The loop stops when no more tweets are returned or when you hit a custom total limit (adjustable if needed).

import datetime as dt
import os
import pandas as pd
import numpy as np  # Add this missing import
import tweepy as tw

# 定义开发者权限
consumer_key = 'xxxxxxxx'
consumer_secret = 'xxxxxxxx'
access_token = 'xxxxxx'
access_token_secret = 'xxxxxxx'

# 访问Twitter API
auth = tw.OAuthHandler(consumer_key, consumer_secret)
auth.set_access_token(access_token, access_token_secret)
api = tw.API(auth, wait_on_rate_limit=True)

# 收集推文的函数 - 支持分页获取更多推文
def get_tweets(handle, max_total_tweets=None):
    all_tweets = []
    # 初始调用,不带max_id
    tweets = api.user_timeline(
        screen_name=handle,
        count=200,
        exclude_replies=True,
        include_rts=False,
        tweet_mode="extended"
    )
    all_tweets.extend(tweets)
    
    # 如果没有推文,直接返回空DataFrame
    if not all_tweets:
        print(f"{handle} has no tweets to extract.\n")
        return pd.DataFrame()
    
    # 循环获取更多推文,直到没有更多或者达到max_total_tweets
    while True:
        # 设置max_id为最后一条推文的ID减1,避免重复获取同一条
        last_tweet_id = all_tweets[-1].id - 1
        
        # 调用API获取下一批推文
        tweets = api.user_timeline(
            screen_name=handle,
            count=200,
            exclude_replies=True,
            include_rts=False,
            tweet_mode="extended",
            max_id=last_tweet_id
        )
        
        # 如果没有更多推文,退出循环
        if not tweets:
            break
        
        all_tweets.extend(tweets)
        
        # 如果设置了max_total_tweets,检查是否达到
        if max_total_tweets and len(all_tweets) >= max_total_tweets:
            all_tweets = all_tweets[:max_total_tweets]
            break
    
    print(f"{handle} Number of tweets extracted: {len(all_tweets)}\n")
    
    # 转换为DataFrame,保持原格式
    df = pd.DataFrame(data=[tweet.user.screen_name for tweet in all_tweets], columns=['handle'])
    df['tweets'] = np.array([tweet.full_text for tweet in all_tweets])
    df['date'] = np.array([tweet.created_at for tweet in all_tweets])
    df['len'] = np.array([len(tweet.full_text) for tweet in all_tweets])
    df['like_count'] = np.array([tweet.favorite_count for tweet in all_tweets])
    df['rt_count'] = np.array([tweet.retweet_count for tweet in all_tweets])
    
    return df

# 所有候选人的Twitter账号列表
handles = ['@JoeBiden', '@ewarren', '@BernieSanders', '@MikeBloomberg', '@PeteButtigieg', '@AndrewYang', '@AmyKlobuchar']
df = pd.DataFrame()

# 遍历不同候选人的账号并收集推文
for handle in handles:
    # 如果你想限制每个账号最多获取N条推文,比如1000,就传max_total_tweets=1000
    df_new = get_tweets(handle)
    df = pd.concat([df, df_new], ignore_index=True)  # 用ignore_index避免索引重复

Key Changes Explained:

  • Pagination Loop: We start with the initial 200 tweets, then keep fetching older batches by setting max_id to the last tweet's ID minus 1 (this ensures we don't re-fetch the final tweet from the previous batch).
  • Stopping Conditions: The loop exits when the API returns no more tweets, or if you set a max_total_tweets limit (useful if you don't want to collect every single tweet from an account).
  • DataFrame Integrity: We build a single list of all tweets first, then convert it to a DataFrame once—this preserves your original column structure and avoids repeated concatenation during the loop.
  • Ignore Index: Added ignore_index=True to pd.concat to prevent duplicate index values in the final combined DataFrame.

How It Works:

When you run this code, Tweepy will automatically handle rate limits by waiting when it hits the 180-call per 15-minute window. It will keep fetching tweets for each handle until there are no older tweets left (or until you hit your max_total_tweets limit).

For example, you might see output like:

@JoeBiden Number of tweets extracted: 3200

@ewarren Number of tweets extracted: 2850

...

内容的提问来源于stack exchange,提问作者Mitchell.Laferla

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.06 15:27:30