如何让WebScrape/Tweet脚本7×24小时循环运行且避免报错?
解决爬虫403拦截与推特重复推文的问题
你遇到的两个错误分别对应路透社反爬拦截和推特禁止重复发布相同内容,下面分点给出具体解决办法:
一、解决403 Forbidden(爬虫被网站拦截)
路透社会识别非浏览器的请求标识,默认requests.get的请求头会被判定为爬虫,需要伪装成浏览器请求:
- 添加
User-Agent请求头,模拟真实浏览器 - 增加随机请求间隔,避免频繁触发反爬机制
- 加入异常处理,防止请求失败导致脚本直接中断
修改后的scrape函数示例:
import requests from bs4 import BeautifulSoup import time import random def scrape(): # 替换成你自己浏览器的User-Agent(可在浏览器F12→网络面板中获取) headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36' } try: # 随机延迟1-3秒,降低请求频率 time.sleep(random.uniform(1, 3)) page = requests.get("https://www.reuters.com/business/future-of-money/", headers=headers) page.raise_for_status() # 主动抛出HTTP请求错误 soup = BeautifulSoup(page.content, "html.parser") home = soup.find(class_="editorial-franchise-layout__main__3cLBl") if not home: print("未找到目标内容区域") return None posts = home.find_all(class_="text__text__1FZLe text__dark-grey__3Ml43 text__inherit-font__1Y8w3 text__inherit-size__1DZJi link__underline_on_hover__2zGL4") if not posts: print("未找到文章列表") return None top_post = posts[0].find("h3", class_="text__text__1FZLe text__dark-grey__3Ml43 text__medium__1kbOh text__heading_3__1kDhc heading__base__2T28j heading__heading_3__3aL54 hero-card__title__33EFM").find_all("span")[0].text.strip() return top_post except Exception as e: print(f"爬取失败:{str(e)}") return None
二、解决187重复推文错误
推特禁止发布完全相同的状态,需要记录已发布内容,仅推送新内容:
- 用本地文件存储已发布的标题,每次爬取后先做重复检查
- 将推特认证逻辑提到全局,避免每次推文重复初始化API
修改后的tweet函数及全局逻辑:
import tweepy # 全局初始化推特API,仅执行一次 api_key = 'deletedforprivacy' api_key_secret = 'deletedforprivacy' access_token = 'deletedforprivacy' access_token_secret = 'deletedforprivacy' authenticator = tweepy.OAuthHandler(api_key, api_key_secret) authenticator.set_access_token(access_token, access_token_secret) api = tweepy.API(authenticator, wait_on_rate_limit=True) def tweet(top_post): # 读取已发布的标题记录 posted_file = "posted_titles.txt" try: with open(posted_file, "r", encoding="utf-8") as f: posted_titles = f.read().splitlines() except FileNotFoundError: posted_titles = [] if top_post in posted_titles: print("该内容已发布过,跳过") return # 发布推文并记录 try: api.update_status(f"{top_post}\nSource : https://www.reuters.com/business/future-of-money/") print(f"推文发布成功:{top_post}") with open(posted_file, "a", encoding="utf-8") as f: f.write(top_post + "\n") except tweepy.TweepyException as e: print(f"推文失败:{str(e)}")
三、循环运行脚本的逻辑
设置固定间隔(比如每小时检查一次),持续监控并推送新内容:
def main(): while True: top_post = scrape() if top_post: tweet(top_post) # 每小时运行一次(3600秒),可根据需求调整间隔 print("等待下一次检查...") time.sleep(3600) if __name__ == "__main__": main()
额外注意事项
User-Agent建议使用你自己常用浏览器的标识,避免被网站识别为通用爬虫- 不要将循环间隔设置过短,防止被路透社拉黑或触发推特限流
- 若路透社页面结构更新,需要同步调整爬虫中的class选择器,否则会无法抓取内容
内容的提问来源于stack exchange,提问作者0xCAJ
相关产品推荐
相关产品推荐

