You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Python借助Tweepy提取推文过程中遇到报错应如何处理?

Tweepy爬取无效推文ID崩溃修复方案

核心问题:api.get_status()请求对应ID的推文时,若目标推文已删除、所属账号封禁/设为私密,会直接抛出异常,未捕获异常会导致程序终止运行。

修复方案

  • 导入Tweepy异常类,在循环中添加异常捕获逻辑,跳过无效ID,可同步记录失败的ID方便后续排查
  • 初始化API时开启自动等待速率限制参数,避免触发API请求频率上限导致崩溃
import pandas as pd
import tweepy
# Tweepy 3.x版本用TweepError,4.x及以上版本替换为TweepyException
from tweepy import TweepError

# 初始化API时建议开启限流自动等待,替换为你自己的认证逻辑
auth = tweepy.OAuthHandler("你的consumer_key", "你的consumer_secret")
auth.set_access_token("你的access_token", "你的access_token_secret")
# Tweepy 4.x版本移除wait_on_rate_limit_notify参数即可
api = tweepy.API(auth, wait_on_rate_limit=True, wait_on_rate_limit_notify=True)

# 创建存储结果的DataFrame
db = pd.DataFrame(columns=['username', 'description', 'location', 'following',
                           'followers', 'totaltweets', 'retweetcount', 'text', 'hashtags'])
# 存储拉取失败的tweet ID,方便后续核对
failed_ids = []

# 从文件读取tweet ID
df = pd.read_excel('dataid.xlsx') 
mylist = df['tweet_id'].tolist()

# 计数器
n=1

# 循环拉取推文
for i in mylist:
    try:
        tweets=api.get_status(i, tweet_mode="extended")
        username = tweets.user.screen_name
        description = tweets.user.description
        location = tweets.user.location
        following = tweets.user.friends_count
        followers = tweets.user.followers_count
        totaltweets = tweets.user.statuses_count
        retweetcount = tweets.retweet_count
        text=tweets.full_text
        # 补充原代码缺失的hashtag提取逻辑,不需要可删除
        hashtext = [hashtag['text'] for hashtag in tweets.entities['hashtags']]
        ith_tweet = [username, description, location, following,followers, totaltweets, retweetcount, text, hashtext]
        db.loc[len(db)] = ith_tweet
        n=n+1
    except TweepError as e:
        print(f"拉取ID {i} 失败,错误信息:{e.reason}")
        failed_ids.append(i)
        continue

# 保存拉取结果
filename = 'scraped_tweets.csv'
db.to_csv(filename, index=False, encoding='utf-8-sig')
# 保存失败的ID列表
pd.DataFrame({'failed_tweet_id': failed_ids}).to_csv('failed_ids.csv', index=False)

优化建议

  • 单条拉取效率较低,可使用api.lookup_statuses()批量拉取推文,单次最多支持查询100个ID,大幅减少请求次数
  • 若数据量较大,建议每拉取100-500条就临时存一次结果,避免程序异常退出导致已拉取的数据全部丢失

内容的提问来源于stack exchange,提问作者Liya

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.27 19:45:07