You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Tweepy爬取Twitter粉丝如何实现断点续爬适配API限流

Twitter粉丝列表断点续爬实现方案

问题核心

Twitter API粉丝接口限流规则如下:

  • 单次连续请求最多返回15-20页数据,单页最大返回200条粉丝记录
  • 触发限流后需要等待15分钟冷却期才可继续发起请求
    原有代码每次重启都会从粉丝列表第一页开始爬取,会重复消耗API配额,以下方案可以实现中断后自动从上次断点位置续爬,不需要手动干预。

实现逻辑

  • 爬取启动时优先加载本地已保存的粉丝数据、以及上次中断时留存的分页游标,直接跳过已爬取的内容
  • 每成功爬取1页数据就立刻持久化到本地,同时更新分页游标断点,避免异常退出时内存数据丢失
  • 明确捕获限流异常,触发限流时自动等待15分钟冷却后继续爬取,不需要手动重启脚本
  • 所有数据爬取完成后自动清理临时断点文件

可直接复用的代码

import time
import os
import pandas as pd
import tweepy
from tweepy import RateLimitError

# 替换为自己的Twitter API认证信息
auth = tweepy.OAuthHandler("API_KEY", "API_SECRET")
auth.set_access_token("ACCESS_TOKEN", "ACCESS_TOKEN_SECRET")
api = tweepy.API(auth, wait_on_rate_limit=False)

def fetch_followers_to_df():
    screen_name = "JeffBezos"
    save_path = "raw_followers.csv"
    breakpoint_path = "follower_breakpoint.txt"
    page_size = 200

    # 初始化数据集:存在历史数据则直接加载,不存在则新建空表
    if os.path.exists(save_path):
        df_followers = pd.read_csv(save_path)
        print(f"加载历史数据,已爬取粉丝数:{len(df_followers)}")
    else:
        df_followers = pd.DataFrame(columns=["User", "ID", "Bio", "Location"])

    # 读取断点分页游标
    next_cursor = -1
    if os.path.exists(breakpoint_path):
        with open(breakpoint_path, "r") as f:
            cursor_val = f.read().strip()
            if cursor_val:
                next_cursor = int(cursor_val)
        print(f"从断点位置恢复爬取,当前游标:{next_cursor}")

    # 获取目标账号基础信息
    user = api.get_user(screen_name=screen_name)
    num_followers = user.followers_count
    total_pages = (num_followers + page_size - 1) // page_size
    print(f"目标账号总粉丝数:{num_followers},预计总页数:{total_pages}")

    # 从断点开始循环爬取
    while True:
        try:
            cursor = tweepy.Cursor(
                api.get_followers,
                screen_name=screen_name,
                count=page_size,
                cursor=next_cursor
            ).pages()

            for page in cursor:
                # 解析当前页粉丝数据
                page_data = []
                for u in page:
                    page_data.append({
                        "User": u.screen_name,
                        "ID": u.id,
                        "Bio": u.description,
                        "Location": u.location
                    })
                # 合并到总数据集
                temp_df = pd.DataFrame(page_data)
                df_followers = pd.concat([df_followers, temp_df], ignore_index=True)
                # 每页爬完立刻写入本地保存
                df_followers.to_csv(save_path, index=False)
                # 更新下一页游标,写入断点文件
                next_cursor = cursor.next_cursor
                with open(breakpoint_path, "w") as f:
                    f.write(str(next_cursor))
                print(f"已完成爬取:{len(df_followers)}/{num_followers} 条粉丝数据")

                # 游标为0代表所有数据爬取完成,退出循环
                if next_cursor == 0:
                    break
            # 爬取完成后删除临时断点文件
            if os.path.exists(breakpoint_path):
                os.remove(breakpoint_path)
            break

        except RateLimitError:
            print("触发API限流,等待15分钟冷却后继续...")
            time.sleep(15*60)
            continue
        except Exception as e:
            print(f"爬取出错:{str(e)},当前进度已保存,下次启动可从断点续爬")
            break

    return df_followers

if __name__ == "__main__":
    df = fetch_followers_to_df()
    print(f"爬取完成,共获取粉丝数据{len(df)}条,已保存到raw_followers.csv")

注意事项

  • 不要手动修改raw_followers.csv和follower_breakpoint.txt两个文件,否则会导致游标错位,出现重复爬取或者漏爬的问题
  • 代码默认关闭了tweepy自带的wait_on_rate_limit参数,自行实现限流等待逻辑,方便自定义等待时的日志输出
  • 单页200条是接口支持的最大值,不要调大该参数,会直接触发接口报错
  • 如果中途需要终止脚本,直接停止即可,当前进度已经实时保存,下次启动会自动从最后一次成功爬取的页面之后继续

内容的提问来源于stack exchange,提问作者amarokWPcom

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.27 02:27:20