You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何修改Tweepy代码避免重复存储含关键词的Twitter推文

解决Twitter推文去重存储的问题

需求

  • 实时从多个Twitter用户中抓取包含指定关键词(比如"Twitter")的最新推文
  • 仅存储未在store中出现过的推文,已存在的直接跳过,持续监测

现有代码问题

原代码每次获取到含关键词的推文就直接存入store,没有做去重判断,会导致重复存储相同推文;同时只支持单个用户监测,不符合多用户需求。

修改后的代码

import tweepy
import time

# 替换为你的API密钥
API_KEY = 'APIKEY'
API_SECRET = 'APISQ'
ACCESS_TOKEN = 'TOK'
ACCESS_TOKEN_SECRET = 'TOkSQ'

auth = tweepy.OAuthHandler(API_KEY, API_SECRET)
auth.set_access_token(ACCESS_TOKEN, ACCESS_TOKEN_SECRET)
api = tweepy.API(auth, wait_on_rate_limit=True)

# 存储推文文本(直接存字符串,方便去重判断)
store = []
# 要监测的多个用户列表
target_users = ['somename', 'anotheruser', 'thirduser']
# 指定关键词
target_keyword = 'Twitter'

while True:
    for username in target_users:
        try:
            # 获取用户最新1条推文
            tweets = api.user_timeline(screen_name=username, count=1, tweet_mode='extended')
            if not tweets:
                continue
            # 取完整推文文本(避免截断)
            tweet_text = tweets[0].full_text
            # 判断:含关键词 且 未在store中
            if target_keyword in tweet_text and tweet_text not in store:
                store.append(tweet_text)
                print(f"新增推文(来自@{username}):{tweet_text}")
                print(f"当前存储的推文总数:{len(store)}")
        except tweepy.TweepyException as e:
            print(f"获取@{username}推文时出错:{str(e)}")
    # 每5秒轮询一次所有用户
    time.sleep(5)

关键改动说明

  • 去重逻辑:把store改为存储推文字符串(而非原代码的嵌套列表),直接用tweet_text not in store判断是否已存在,避免重复存储
  • 多用户支持:将单个用户名改为用户列表target_users,通过循环遍历实现多用户监测
  • 完整推文获取:添加tweet_mode='extended'参数,避免推文内容被截断
  • 异常处理:捕获Tweepy请求异常,防止单个用户请求失败导致整个程序崩溃
  • 优化输出:添加明确的新增提示,方便查看监测状态

内容的提问来源于stack exchange,提问作者Snuskbetty

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.11 11:55:20