You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将Flask网页与爬虫计数器关联,实时显示当前爬取页码

实现Flask爬虫页码计数器实时前端展示

核心思路

通过后端暴露计数器API + 前端定时轮询的方式,将爬虫的页码计数器实时同步到页面。以下是具体实现步骤:


1. 改造爬虫模块(scraper.py)

将计数器改为线程安全的全局变量,避免多线程爬取时出现竞态问题:

from threading import Lock

# 全局爬取页码计数器,初始为0
crawl_counter = 0
# 线程锁,保证计数器读写的原子性
counter_lock = Lock()

def crawl_single_page(page_num):
    """爬取单页数据的函数"""
    global crawl_counter
    # 更新计数器时加锁,防止多线程冲突
    with counter_lock:
        crawl_counter = page_num
    
    # 你的爬取逻辑(请求第三方网站、解析数据等)
    # ...

def start_crawling(max_page):
    """启动批量爬取的入口函数"""
    for page in range(1, max_page + 1):
        crawl_single_page(page)
        # 模拟爬取延迟,实际根据网站响应速度调整
        # time.sleep(1)

2. 改造Flask后端主文件

添加两个关键路由:一个用于启动爬虫(非阻塞),另一个用于获取当前计数器值:

from flask import Flask, render_template, jsonify
from scraper import crawl_counter, counter_lock, start_crawling
import threading

app = Flask(__name__)

# 渲染爬取页面
@app.route('/')
def scrape_page():
    return render_template('scrape.html')

# 提供计数器查询API
@app.route('/api/crawl-counter')
def get_crawl_counter():
    with counter_lock:
        current_page = crawl_counter
    return jsonify({"current_page": current_page})

# 启动爬虫的接口(非阻塞)
@app.route('/api/start-crawl')
def trigger_crawl():
    # 用线程启动爬虫,避免阻塞Flask主线程
    threading.Thread(target=start_crawling, args=(100,), daemon=True).start()
    return jsonify({"status": "爬取已启动"})

if __name__ == '__main__':
    app.run(debug=True)

3. 改造前端页面(scrape.html)

在两个按钮之间添加计数器展示元素,并通过JavaScript定时轮询后端API更新内容:

<!DOCTYPE html>
<html lang="zh-CN">
<head>
    <meta charset="UTF-8">
    <title>爬虫状态监控</title>
</head>
<body>
    <button onclick="startCrawl()">开始爬取</button>
    <!-- 计数器展示区域 -->
    <span id="counter-display">当前爬取页码:0</span>
    <button onclick="stopCrawl()">停止爬取</button>

    <script>
        // 每秒轮询一次计数器API
        setInterval(() => {
            fetch('/api/crawl-counter')
                .then(res => res.json())
                .then(data => {
                    document.getElementById('counter-display').textContent = `当前爬取页码:${data.current_page}`;
                })
                .catch(err => console.error('获取计数器失败:', err));
        }, 1000);

        // 触发爬虫启动
        function startCrawl() {
            fetch('/api/start-crawl')
                .then(res => res.json())
                .then(data => alert(data.status));
        }

        // 停止爬取的逻辑(需在爬虫中添加停止标志)
        function stopCrawl() {
            // 示例:调用停止API(需后端实现对应的路由和停止逻辑)
            fetch('/api/stop-crawl')
                .then(res => res.json())
                .then(data => alert(data.status));
        }
    </script>
</body>
</html>

进阶优化(可选)

如果需要更实时的更新(避免轮询的延迟和资源消耗),可以用WebSocket替代轮询,比如使用flask-socketio:

  1. 安装依赖:pip install flask-socketio
  2. 后端改造:在爬虫更新计数器时主动向前端推送事件
  3. 前端监听WebSocket事件,实时更新页面

内容的提问来源于stack exchange,提问作者SolidOpt

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.27 05:12:39