You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python多次运行Tweepy抓取函数如何不覆盖原有推文IPM采集数据

解决方案

核心逻辑是通过**推文唯一ID(tweet_id)**作为匹配键,维护历史采集结果,将每次新的IPM值追加到对应推文的IPM列表中,具体实现如下:

1. 核心调整点

  • 新增tweet_id作为唯一主键,用于匹配多次运行时的同一条推文,解决新数据被识别为新推文的问题
  • 将原ipm字段改为ipm_history列表类型,存储每次采集的IPM值
  • 可选新增collect_time列表,记录每次IPM对应的采集时间,方便后续分析变化趋势
  • 函数支持传入历史采集结果作为参数,更新完成后返回新的历史数据集

2. 修改后代码

import tweepy
import pandas as pd
import traceback
import datetime

def update_tweet_data(username: str, count: int, history_df: pd.DataFrame = None) -> pd.DataFrame:
    # 首次运行时初始化历史DataFrame
    if history_df is None:
        history_df = pd.DataFrame(columns=[
            'tweet_id', 'retweets', 'favorites', 'total_interactions', 'ipm_history', 'collect_time'
        ])
    
    try:
        for tweet in tweepy.Cursor(api.user_timeline, id=username).items(count):
            tweet_id = str(tweet.id)
            current_collect_time = datetime.datetime.utcnow()
            timesince = current_collect_time - tweet.created_at
            ipm = (tweet.favorite_count + tweet.retweet_count) / (abs(int(timesince.total_seconds() / 60)))
            
            # 匹配已有推文更新数据
            if tweet_id in history_df['tweet_id'].values:
                idx = history_df[history_df['tweet_id'] == tweet_id].index[0]
                # 更新最新的互动数
                history_df.at[idx, 'retweets'] = tweet.retweet_count
                history_df.at[idx, 'favorites'] = tweet.favorite_count
                history_df.at[idx, 'total_interactions'] = tweet.favorite_count + tweet.retweet_count
                # 追加IPM和采集时间
                history_df.at[idx, 'ipm_history'].append(ipm)
                history_df.at[idx, 'collect_time'].append(current_collect_time)
            # 新增未收录的推文
            else:
                new_row = {
                    'tweet_id': tweet_id,
                    'retweets': tweet.retweet_count,
                    'favorites': tweet.favorite_count,
                    'total_interactions': tweet.favorite_count + tweet.retweet_count,
                    'ipm_history': [ipm],
                    'collect_time': [current_collect_time]
                }
                history_df = pd.concat([history_df, pd.DataFrame([new_row])], ignore_index=True)
        
        return history_df
    except tweepy.RateLimitError:
        print("Rate limit exceeded")
        return history_df
    except tweepy.TweepError as err:
        print(f"Error: {err.reason}")
        return history_df
    except Exception:
        traceback.print_exc()
        return history_df

3. 使用方式

多次调用函数时传入上一次返回的历史DataFrame即可实现数据持续更新:

# 首次运行采集基础数据
history = update_tweet_data("目标用户名", 10)
# 间隔一段时间后再次运行,更新IPM历史
history = update_tweet_data("目标用户名", 10, history_df=history)
# 后续每次运行都传入上一次的history变量即可

如果需要持久化存储数据,可以在每次更新后将历史数据保存为pickle格式:

# 保存
history.to_pickle("tweet_ipm_history.pkl")
# 下次运行读取
history = pd.read_pickle("tweet_ipm_history.pkl")

内容的提问来源于stack exchange,提问作者sak

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.02 17:45:03