如何将Flask网页与爬虫计数器关联,实时显示当前爬取页码
实现Flask爬虫页码计数器实时前端展示
核心思路
通过后端暴露计数器API + 前端定时轮询的方式,将爬虫的页码计数器实时同步到页面。以下是具体实现步骤:
1. 改造爬虫模块(scraper.py)
将计数器改为线程安全的全局变量,避免多线程爬取时出现竞态问题:
from threading import Lock # 全局爬取页码计数器,初始为0 crawl_counter = 0 # 线程锁,保证计数器读写的原子性 counter_lock = Lock() def crawl_single_page(page_num): """爬取单页数据的函数""" global crawl_counter # 更新计数器时加锁,防止多线程冲突 with counter_lock: crawl_counter = page_num # 你的爬取逻辑(请求第三方网站、解析数据等) # ... def start_crawling(max_page): """启动批量爬取的入口函数""" for page in range(1, max_page + 1): crawl_single_page(page) # 模拟爬取延迟,实际根据网站响应速度调整 # time.sleep(1)
2. 改造Flask后端主文件
添加两个关键路由:一个用于启动爬虫(非阻塞),另一个用于获取当前计数器值:
from flask import Flask, render_template, jsonify from scraper import crawl_counter, counter_lock, start_crawling import threading app = Flask(__name__) # 渲染爬取页面 @app.route('/') def scrape_page(): return render_template('scrape.html') # 提供计数器查询API @app.route('/api/crawl-counter') def get_crawl_counter(): with counter_lock: current_page = crawl_counter return jsonify({"current_page": current_page}) # 启动爬虫的接口(非阻塞) @app.route('/api/start-crawl') def trigger_crawl(): # 用线程启动爬虫,避免阻塞Flask主线程 threading.Thread(target=start_crawling, args=(100,), daemon=True).start() return jsonify({"status": "爬取已启动"}) if __name__ == '__main__': app.run(debug=True)
3. 改造前端页面(scrape.html)
在两个按钮之间添加计数器展示元素,并通过JavaScript定时轮询后端API更新内容:
<!DOCTYPE html> <html lang="zh-CN"> <head> <meta charset="UTF-8"> <title>爬虫状态监控</title> </head> <body> <button onclick="startCrawl()">开始爬取</button> <!-- 计数器展示区域 --> <span id="counter-display">当前爬取页码:0</span> <button onclick="stopCrawl()">停止爬取</button> <script> // 每秒轮询一次计数器API setInterval(() => { fetch('/api/crawl-counter') .then(res => res.json()) .then(data => { document.getElementById('counter-display').textContent = `当前爬取页码:${data.current_page}`; }) .catch(err => console.error('获取计数器失败:', err)); }, 1000); // 触发爬虫启动 function startCrawl() { fetch('/api/start-crawl') .then(res => res.json()) .then(data => alert(data.status)); } // 停止爬取的逻辑(需在爬虫中添加停止标志) function stopCrawl() { // 示例:调用停止API(需后端实现对应的路由和停止逻辑) fetch('/api/stop-crawl') .then(res => res.json()) .then(data => alert(data.status)); } </script> </body> </html>
进阶优化(可选)
如果需要更实时的更新(避免轮询的延迟和资源消耗),可以用WebSocket替代轮询,比如使用flask-socketio:
- 安装依赖:
pip install flask-socketio - 后端改造:在爬虫更新计数器时主动向前端推送事件
- 前端监听WebSocket事件,实时更新页面
内容的提问来源于stack exchange,提问作者SolidOpt
相关产品推荐
相关产品推荐

