Python调度器异常:无法每晚定时执行网页抓取函数
核心问题
我在树莓派上用schedule库写程序,想让它每晚00:10自动执行网页抓取的scrape函数。换成简单的testTime函数时调度能正常运行,但用scrape函数就失效——只有重启程序时,抓取的数据能存进fixed_scrape变量,之后每晚的自动调度完全没反应。试了Ischedule也没用,作为Python新手,想知道问题出在哪?
原代码
from multiprocessing.sharedctypes import Value from optparse import Values import re, requests, ast from prompt_toolkit import print_formatted_text import RPi.GPIO as GPIO import time from heapq import nsmallest import schedule import pip._vendor.requests HoursPerDay = 10 PriceLimit = 80 GPIO.setmode(GPIO.BCM) GPIO.setwarnings(False) GPIO.setup(18, GPIO.OUT, initial=GPIO.LOW) def scrape(): r = requests.get('https://www.elbruk.se/timpriser-se3-stockholm') return dict(zip([i[0] for i in re.findall(r"'((2[0-4]|[01]?[0-9]):([0-5]?[0-9]))'", r.text)], ast.literal_eval(re.search(r"data: .*(\[.*?\])[\s\S]+(?='Idag snitt')", r.text).group(1)))) fixed_scrape = scrape() def cheap(): my_dict = fixed_scrape cheapest_hour = nsmallest(HoursPerDay, my_dict, key=my_dict.get) return [(i.removesuffix(':00')).lstrip('0') or 0 for i in cheapest_hour] def threshold(): price_hour = fixed_scrape price_hour = {k:v for k,v in price_hour.items() if v < PriceLimit} dick = nsmallest(24, price_hour, key=price_hour.get) return ([(i.removesuffix(':00')).lstrip('0') or 0 for i in dick]) schedule.every().day.at("00:10").do(scrape) while True: schedule.run_pending() now_time = time.strftime("%H") cheapest_hours = cheap() cheap_threshold = threshold() print(cheapest_hours, "Are the cheapest hours today") print(cheap_threshold, "Hours when price is under" , PriceLimit, "öre/kwh. (Buying on threshold)") print("Active hour now is: ",now_time) if now_time in cheap_threshold: GPIO.output(18, GPIO.HIGH) elif now_time in cheapest_hours: GPIO.output(18, GPIO.HIGH) else: GPIO.output(18, GPIO.LOW) time.sleep(30)
问题根源
调度执行但未更新数据:
你调用schedule.every().day.at("00:10").do(scrape)只是让调度执行scrape()函数,但返回的新数据并没有赋值给fixed_scrape。fixed_scrape只在程序启动时被赋值一次,后续调度跑scrape的结果直接被丢弃,cheap()和threshold()一直用的是启动时的旧数据。scrape函数无异常处理,触发异常后调度终止:
网页请求(比如网络断了)或正则解析失败(比如网页结构变了)会抛出错误,schedule默认会吞掉这些异常,导致后续的调度任务不再执行。这就是为什么换成简单的testTime函数没问题,因为它不会抛出异常。time.sleep位置错误:
现在time.sleep(30)写在while循环外面,程序第一次执行完循环体就会卡住,不会再回来检查调度任务。必须把sleep放到循环内部,才能让程序持续循环、定期检查调度。
修复方案
1. 让调度更新fixed_scrape
把调度任务改成一个包装函数,让它把scrape的结果赋值给fixed_scrape:
def update_scrape(): global fixed_scrape # 声明使用全局变量 fixed_scrape = scrape() # 替换原来的调度语句 schedule.every().day.at("00:10").do(update_scrape)
2. 给scrape函数加异常处理
避免请求或解析失败导致调度中断,失败时返回旧数据保证程序继续运行:
def scrape(): try: # 加超时防止请求卡住 r = requests.get('https://www.elbruk.se/timpriser-se3-stockholm', timeout=10) r.raise_for_status() # 触发HTTP错误(比如404、500) # 检查正则匹配结果 time_matches = re.findall(r"'((2[0-4]|[01]?[0-9]):([0-5]?[0-9]))'", r.text) price_match = re.search(r"data: .*(\[.*?\])[\s\S]+(?='Idag snitt')", r.text) if not time_matches or not price_match: raise ValueError("网页结构变化,找不到目标数据") prices = ast.literal_eval(price_match.group(1)) return dict(zip([i[0] for i in time_matches], prices)) except Exception as e: print(f"抓取失败: {str(e)}") return fixed_scrape # 失败时返回旧数据,不影响后续逻辑
3. 调整time.sleep的位置
把time.sleep(30)放到while循环的最后,确保程序能持续循环检查调度:
while True: schedule.run_pending() now_time = time.strftime("%H") cheapest_hours = cheap() cheap_threshold = threshold() print(cheapest_hours, "是今天最便宜的时段") print(cheap_threshold, "电价低于", PriceLimit, "öre/kwh的时段(阈值模式)") print("当前时段: ", now_time) if now_time in cheap_threshold or now_time in cheapest_hours: GPIO.output(18, GPIO.HIGH) else: GPIO.output(18, GPIO.LOW) time.sleep(30) # 放到循环最后
内容的提问来源于stack exchange,提问作者TheSwedishFlu

