Scrapy爬虫自动重试与定时执行的正确循环结构问题
解决方案
1. 用子进程启动Scrapy,解决循环重启问题
直接通过Python的subprocess模块调用Scrapy命令,每次启动都是独立进程,不会残留之前的Scrapy运行环境,完美解决while循环无法重启的问题:
import subprocess import os def run_scrapy_spider(spider_name): # 调用Scrapy命令,指定输出文件(假设你的爬虫名就是home/away/today) result = subprocess.run( ["scrapy", "crawl", spider_name, "-o", f"{spider_name}.json"], capture_output=True, text=True, cwd="/path/to/your/scrapy/project" # 替换成你的Scrapy项目根目录 ) # 返回爬取是否成功(returncode为0表示正常结束) return result.returncode == 0
2. 验证爬取结果完整性
合并数据前先检查三个JSON文件是否有效,避免漏爬导致data()函数出错:
import json def validate_json_file(file_path): if not os.path.exists(file_path): return False try: with open(file_path, 'r', encoding='utf-8') as f: data = json.load(f) # 根据你的数据结构调整验证逻辑,比如检查是否有有效赛事数据 if len(data) == 0: return False # 示例:检查每条数据是否包含赛事ID和比分字段 for item in data: if 'match_id' not in item or 'score' not in item: return False return True except Exception as e: print(f"验证文件{file_path}失败: {str(e)}") return False def check_all_files(): return all(validate_json_file(f) for f in ['home.json', 'away.json', 'today.json'])
3. 带重试的爬取循环
设置最大重试次数,直到爬取完整或达到上限,重试前清理无效文件避免干扰:
def crawl_with_retry(max_retries=3): retry_count = 0 while retry_count < max_retries: print(f"第{retry_count+1}次爬取...") # 依次启动三个爬虫 home_ok = run_scrapy_spider('home') away_ok = run_scrapy_spider('away') today_ok = run_scrapy_spider('today') if home_ok and away_ok and today_ok and check_all_files(): print("爬取完成,开始合并数据") data() # 调用你的合并函数 return True else: print("爬取不完整,准备重试") retry_count += 1 # 删除无效文件,避免下一次爬取读取旧数据 for f in ['home.json', 'away.json', 'today.json']: if os.path.exists(f): os.remove(f) print(f"重试{max_retries}次仍失败,停止爬取") return False
4. 每天定时执行
用schedule库实现凌晨12:05自动执行,先通过pip install schedule安装依赖:
import schedule import time def daily_task(): print("启动每日爬取任务...") crawl_with_retry() # 设置每天00:05执行任务 schedule.every().day.at("00:05").do(daily_task) # 保持程序后台运行 if __name__ == "__main__": print("定时任务已启动,等待执行") while True: schedule.run_pending() time.sleep(60) # 每分钟检查一次任务状态
关键注意点
- 替换
run_scrapy_spider里的cwd参数为你的Scrapy项目实际路径 - 根据实际赛事数据结构修改
validate_json_file里的验证规则 - 可以添加日志模块(比如
logging)记录爬取细节,方便排查问题
内容的提问来源于stack exchange,提问作者user19925544
相关产品推荐
相关产品推荐

