如何避免Heroku Dyno休眠以持续运行Python网络爬虫
我编写了一款Python网络爬虫脚本并部署在Heroku平台,该脚本为持续运行类型,每休眠1分钟后执行爬取操作。但运行数分钟后便停止工作,推测是Heroku Dyno进入休眠状态。请问如何防止Dyno闲置,保障脚本持续运行?是否存在其他故障原因?
附脚本代码:
import time from bs4 import BeautifulSoup import urllib.request import schedule from bs4.element import Tag url = 'url_here' prev_news = [] # Store the previous state of news to compare changes prev_updated = "Initial Value" updated_news = [] def compare_variables(prev, curr): """ Compare two variables and print the comparison result. """ if prev == curr: print("No change values") return False else: print("Change detected") return True def fetch_last_updated() -> str: page = urllib.request.urlopen(url) soup = BeautifulSoup(page.read(), 'html.parser') """ ....... """ return last_updated def fetch_latest_news() -> list: """ ... """ return latest_news def value(component: Tag) -> list: result = [] """ ... """ return result def change_link(url: str) -> str: if url != None: """ ... """ return url else: return None def job(): # Fetch last updated information from DTU Official Webpage curr_updated = fetch_last_updated() print("Runnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnning") # Compare the two variables if compare_variables(prev_updated, curr_updated): # Changes detected in last updated information print("Last Updated has changed to :", curr_updated) curr_news = fetch_latest_news() # Update the previous value for the next iteration prev_updated = curr_updated schedule.every(1).minutes.do(job) if __name__ == "__main__": while True: schedule.run_pending() time.sleep(1)
一、防止Heroku Dyno休眠的方法
Heroku免费Dyno会在闲置30分钟后自动休眠,且每月有550小时的运行时长限制。要持续运行脚本,可参考以下方案:
1. 升级至付费Dyno
选择Hobby或Professional级别的付费Dyno,这类Dyno不会因闲置休眠,且提供更稳定的资源(如更高内存、无时长限制),适合需要持续运行的任务。
2. 为脚本添加Web端点(免费Dyno适配)
免费Dyno仅会在收到Web请求时保持活跃,因此可以给脚本添加一个简单的Web服务,同时通过定时请求该端点来防止休眠:
- 安装Flask库:
pip install flask - 修改脚本,同时启动Web服务和爬虫任务:
# 新增Flask相关代码 import os from flask import Flask import threading app = Flask(__name__) @app.route('/') def hello(): return "Crawler is running!" def run_web_server(): app.run(host='0.0.0.0', port=int(os.environ.get('PORT', 5000))) # 原爬虫代码保持不变... if __name__ == "__main__": # 启动Web服务线程 web_thread = threading.Thread(target=run_web_server) web_thread.daemon = True web_thread.start() # 原爬虫循环 while True: schedule.run_pending() time.sleep(1)
- 部署后,每隔25分钟左右向你的Heroku应用URL发送一次请求(可通过本地定时脚本实现),确保Dyno不会进入休眠。
二、其他可能的故障原因
除了Dyno休眠,脚本停止运行还可能由以下问题导致:
1. 未捕获的异常
你的脚本中urllib.request.urlopen未处理网络异常(如连接超时、目标网站拒绝访问),fetch_last_updated等函数内部若存在未定义变量(如省略部分中的last_updated未正确赋值),都会导致脚本崩溃退出。建议在关键代码块添加异常捕获:
def fetch_last_updated() -> str: try: page = urllib.request.urlopen(url, timeout=10) soup = BeautifulSoup(page.read(), 'html.parser') # 省略的解析逻辑... return last_updated except Exception as e: print(f"Fetch error: {e}") return prev_updated # 返回上一次值避免逻辑中断
2. 全局变量修改错误
job函数中修改prev_updated时未声明global,会创建局部变量,导致全局的prev_updated始终不变,可能引发逻辑错误(虽不会直接导致停止,但可能间接引发其他问题)。需修改函数:
def job(): global prev_updated # 声明使用全局变量 curr_updated = fetch_last_updated() # 后续逻辑... prev_updated = curr_updated
3. 资源限制触发终止
免费Dyno仅提供512MB内存,若爬虫存在内存泄漏(如未及时释放大对象),或单次爬取消耗过多资源,会被Heroku强制终止。可通过heroku logs --tail查看实时日志,排查是否有内存溢出或进程被杀死的记录。
4. 目标网站反爬机制
Heroku的IP可能被目标网站屏蔽,导致请求失败。可尝试添加请求头模拟浏览器访问:
def fetch_last_updated() -> str: headers = {'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/91.0.4472.124 Safari/537.36'} req = urllib.request.Request(url, headers=headers) page = urllib.request.urlopen(req) # 后续逻辑...
内容的提问来源于stack exchange,提问作者Navi_Here

