You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Flask脚本中定时重新抓取URL数据并更新页面?

解决方案:用APScheduler实现Flask页面定时更新数据

疑问解答

  • 是否用APScheduler? 必须用,这是最适合你场景的方案:脚本内部就能实现定时任务,不用依赖外部工具,更新后的数据直接存在内存里,用户刷新页面就能看到最新内容。
  • Windows任务计划? 不推荐,每次更新都要终止旧脚本再启动新的,还要处理进程管理,操作繁琐,不如APScheduler简洁。
  • 脚本内部实现定期更新? 完全可行,下面给你基于原代码的修改版本。

问题根源

你之前的代码里,pictures、titles、current_time都是启动时初始化的全局变量,APScheduler的sensor函数里的变量是局部的,根本没修改全局变量,所以页面一直显示旧数据。

具体实现代码

把抓取逻辑封装成函数,用APScheduler定时调用它更新全局变量,同时处理Flask debug模式的重复启动问题:

# News feed test for Xibo Signage
from flask import Flask, render_template
from bs4 import BeautifulSoup
import requests
from datetime import datetime
from apscheduler.schedulers.background import BackgroundScheduler
import os

app = Flask(__name__)

# 全局变量,存储抓取的数据和更新时间
pictures = []
titles = []
current_time = datetime.now().strftime("%d/%m/%Y %H:%M:%S")

def fetch_data():
    """定时抓取数据并更新全局变量"""
    global pictures, titles, current_time
    url = "https://news.clemson.edu/tag/extension/"
    try:
        # 发起请求,模拟浏览器UA
        response = requests.get(url, headers={'user-agent':'Mozilla/5.0'})
        response.raise_for_status()  # 捕获请求异常
        soup = BeautifulSoup(response.text, 'html.parser')
        
        # 抓取图片
        new_pictures = []
        for e in soup.select('article img.lazyload'):
            src = e.get('data-src')
            if src:  # 过滤空链接
                new_pictures.append(src)
        
        # 抓取标题
        new_titles = []
        for e in soup.select('article header'):
            title_tag = e.find("h3", class_="entry-title bold")
            if title_tag:  # 过滤找不到标题的情况
                new_titles.append(title_tag.text.strip())
        
        # 只有抓取成功才更新全局变量
        pictures = new_pictures
        titles = new_titles
        # 更新当前时间
        current_time = datetime.now().strftime("%d/%m/%Y %H:%M:%S")
        print(f"数据更新完成,时间:{current_time}")
    except Exception as e:
        print(f"抓取失败:{str(e)},保留旧数据")

@app.route('/') 
def home():
    # 每次请求页面时,直接返回最新的全局变量
    return render_template('home.html', pictures=pictures, titles=titles, current_time=current_time)

if __name__ == '__main__':
    # 初始化调度器,每小时执行一次fetch_data
    scheduler = BackgroundScheduler(daemon=True)
    scheduler.add_job(fetch_data, 'interval', hours=1)
    # 启动时先抓取一次数据
    fetch_data()
    
    # 处理Flask debug模式下重复启动调度器的问题
    if not app.debug or os.environ.get('WERKZEUG_RUN_MAIN') == 'true':
        scheduler.start()
    
    # 启动Flask服务
    app.run(host='0.0.0.0', debug=True)

代码说明

  1. 全局变量:把pictures、titles、current_time声明为全局变量,方便定时函数修改。
  2. fetch_data函数:封装所有抓取逻辑,异常处理确保抓取失败时不会清空旧数据,只会打印错误信息。
  3. APScheduler配置:用BackgroundScheduler后台运行,每小时执行一次fetch_data,启动时先手动调用一次初始化数据。
  4. Debug模式处理:Flask debug模式下会自动重启进程,加判断避免调度器被初始化多次。

如果非要用Windows任务计划(不推荐)

  1. 创建任务:打开「任务计划程序」,创建基本任务,设置触发频率为每小时。
  2. 终止旧脚本:在启动新脚本前,先执行终止命令。为了精准杀死当前脚本进程,建议在脚本启动时记录PID:
    • 在脚本开头添加:
      import os
      with open('flask_app.pid', 'w') as f:
          f.write(str(os.getpid()))
      
    • 任务计划的操作里,先执行:cmd /c taskkill /f /pid $(type flask_app.pid)(或者写个bat脚本处理),再启动虚拟环境里的python脚本。
  3. 启动脚本:设置操作为启动程序,路径选虚拟环境的python.exe,参数选你的脚本路径。

内容的提问来源于stack exchange,提问作者Total30

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.28 04:57:31