Python多次运行Tweepy抓取函数如何不覆盖原有推文IPM采集数据
解决方案
核心逻辑是通过**推文唯一ID(tweet_id)**作为匹配键,维护历史采集结果,将每次新的IPM值追加到对应推文的IPM列表中,具体实现如下:
1. 核心调整点
- 新增
tweet_id作为唯一主键,用于匹配多次运行时的同一条推文,解决新数据被识别为新推文的问题 - 将原
ipm字段改为ipm_history列表类型,存储每次采集的IPM值 - 可选新增
collect_time列表,记录每次IPM对应的采集时间,方便后续分析变化趋势 - 函数支持传入历史采集结果作为参数,更新完成后返回新的历史数据集
2. 修改后代码
import tweepy import pandas as pd import traceback import datetime def update_tweet_data(username: str, count: int, history_df: pd.DataFrame = None) -> pd.DataFrame: # 首次运行时初始化历史DataFrame if history_df is None: history_df = pd.DataFrame(columns=[ 'tweet_id', 'retweets', 'favorites', 'total_interactions', 'ipm_history', 'collect_time' ]) try: for tweet in tweepy.Cursor(api.user_timeline, id=username).items(count): tweet_id = str(tweet.id) current_collect_time = datetime.datetime.utcnow() timesince = current_collect_time - tweet.created_at ipm = (tweet.favorite_count + tweet.retweet_count) / (abs(int(timesince.total_seconds() / 60))) # 匹配已有推文更新数据 if tweet_id in history_df['tweet_id'].values: idx = history_df[history_df['tweet_id'] == tweet_id].index[0] # 更新最新的互动数 history_df.at[idx, 'retweets'] = tweet.retweet_count history_df.at[idx, 'favorites'] = tweet.favorite_count history_df.at[idx, 'total_interactions'] = tweet.favorite_count + tweet.retweet_count # 追加IPM和采集时间 history_df.at[idx, 'ipm_history'].append(ipm) history_df.at[idx, 'collect_time'].append(current_collect_time) # 新增未收录的推文 else: new_row = { 'tweet_id': tweet_id, 'retweets': tweet.retweet_count, 'favorites': tweet.favorite_count, 'total_interactions': tweet.favorite_count + tweet.retweet_count, 'ipm_history': [ipm], 'collect_time': [current_collect_time] } history_df = pd.concat([history_df, pd.DataFrame([new_row])], ignore_index=True) return history_df except tweepy.RateLimitError: print("Rate limit exceeded") return history_df except tweepy.TweepError as err: print(f"Error: {err.reason}") return history_df except Exception: traceback.print_exc() return history_df
3. 使用方式
多次调用函数时传入上一次返回的历史DataFrame即可实现数据持续更新:
# 首次运行采集基础数据 history = update_tweet_data("目标用户名", 10) # 间隔一段时间后再次运行,更新IPM历史 history = update_tweet_data("目标用户名", 10, history_df=history) # 后续每次运行都传入上一次的history变量即可
如果需要持久化存储数据,可以在每次更新后将历史数据保存为pickle格式:
# 保存 history.to_pickle("tweet_ipm_history.pkl") # 下次运行读取 history = pd.read_pickle("tweet_ipm_history.pkl")
内容的提问来源于stack exchange,提问作者sak
相关产品推荐
相关产品推荐

