You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Tweepy批量抓取用户时间线报错及相关优化问题咨询

解决方案:批量抓取Twitter推文的常见问题处理

Hey there! Let's tackle your three core issues step by step—since you're working on an economics thesis, getting this workflow smooth will let you focus on the analysis instead of debugging.

1. 循环抓取粉丝推文时触发AttributeError

This usually happens when one of the user objects in your follower list is incomplete, deleted, or suspended (so Tweepy can't retrieve all its attributes like id or screen_name). Here's how to fix it:

排查与修复步骤:

  • 添加错误处理跳过异常用户:把抓取代码包裹在try-except块中,捕获错误并记录问题用户,然后继续处理下一个用户。
  • 验证粉丝列表完整性:Twitter限制get_followers每次请求最多返回200个用户,如果粉丝数量超过200,需要用cursor参数分页获取完整列表。

示例代码片段:

import tweepy

# 初始化API并开启速率限制自动等待(后面会详细讲)
auth = tweepy.OAuthHandler("CONSUMER_KEY", "CONSUMER_SECRET")
auth.set_access_token("ACCESS_TOKEN", "ACCESS_TOKEN_SECRET")
api = tweepy.API(auth, wait_on_rate_limit=True)

# 分页获取完整粉丝列表
def get_all_followers(user_id):
    followers = []
    for page in tweepy.Cursor(api.get_followers, user_id=user_id, count=200).pages():
        followers.extend(page)
    return followers

follower_list = get_all_followers("目标网红的用户ID")

# 带错误处理的循环抓取
for follower in follower_list:
    try:
        # 抓取该粉丝的指定日期推文
        tweets = api.user_timeline(
            user_id=follower.id,
            count=200,
            tweet_mode="extended",
            since="2022-01-01",
            until="2023-01-01"
        )
        # 处理推文(保存到CSV/数据库等)
        for tweet in tweets:
            print(f"{follower.screen_name}: {tweet.full_text[:100]}")
    except AttributeError as e:
        # 记录问题用户并继续
        user_identifier = follower.screen_name if hasattr(follower, "screen_name") else str(follower.id)
        print(f"跳过用户 {user_identifier}: AttributeError - {str(e)}")
        continue

2. Tweepy按日期筛选推文结果异常

日期筛选在Tweepy v1.1和v2版本中的表现不同,且免费API存在推文数量限制(最多获取用户最近3200条推文)。以下是正确的用法:

针对Tweepy v1.1(使用API类):

  • 使用since(格式YYYY-MM-DD,包含该日期)和until(格式YYYY-MM-DD,不包含该日期)参数。
  • 注意:这两个参数仅过滤最近3200条推文,如果用户在你的日期范围内有更早的推文且超出3200条限制,免费API无法获取。
  • 需要用max_id参数分页获取更早的批次推文。

针对Tweepy v2(使用Client类,推荐新项目使用):

  • 使用start_time和end_time参数,格式为ISO 8601(例如2022-01-01T00:00:00Z)。
  • 该版本筛选更精准,但免费用户仍有推文数量限制。

示例v2代码片段:

client = tweepy.Client(bearer_token="你的Bearer Token", wait_on_rate_limit=True)

def get_tweets_in_date_range(user_id, start_time, end_time):
    tweets = []
    next_token = None
    while True:
        response = client.get_users_tweets(
            id=user_id,
            max_results=100,
            start_time=start_time,
            end_time=end_time,
            tweet_fields=["created_at", "text"],
            pagination_token=next_token
        )
        if not response.data:
            break
        tweets.extend(response.data)
        next_token = response.meta.get("next_token")
        if not next_token:
            break
    return tweets

# 调用示例
tweets = get_tweets_in_date_range(
    follower.id,
    start_time="2022-01-01T00:00:00Z",
    end_time="2023-01-01T00:00:00Z"
)

3. 避免触发Twitter速率限制

Tweepy有内置工具处理速率限制,再配合一些最佳实践就能顺利批量抓取:

  • 开启自动速率限制等待:初始化API或Client时添加wait_on_rate_limit=True参数,Tweepy会自动在速率限制重置时暂停并恢复请求(免费API通常每15分钟重置一次)。
  • 批量请求:使用api.lookup_users()这类方法批量获取用户信息,替代单独的api.get_user()请求,减少调用次数。
  • 记录进度:把已经处理过的用户ID保存到文件(如CSV或JSON),如果脚本崩溃,可以从断点继续,避免重复请求。
  • 避免过度频繁请求:如果处理超大规模粉丝列表,可在循环中加入短暂延迟(比如time.sleep(1)),不过内置的速率限制处理已经足够应对大部分场景。

内容的提问来源于stack exchange,提问作者Carl Wright

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.11 07:41:47