You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于Flask实现HTML动态变量实时更新:文章爬取计数实时展示

Real-Time Article Count for Your Web Scraping Project

Hey there! Let's get that real-time article count working smoothly for your project. Here's a step-by-step breakdown tailored to your setup:

1. Backend: Track Count & Expose via API

First, we need to fix how you track article IDs and counts (using your database instead of a local variable) and create a simple API endpoint so the frontend can fetch the latest number.

Fix Article ID Logic

Instead of manually incrementing an id variable, let your database handle auto-incrementing IDs—most databases like SQLite, PostgreSQL, or MySQL support this out of the box. This ensures IDs stay consistent even if your scraper restarts.

For example, an SQLite articles table might look like this:

CREATE TABLE articles (
    id INTEGER PRIMARY KEY AUTOINCREMENT,
    title TEXT,
    content TEXT,
    -- add other columns you need (like scrape date)
);

Add an API Endpoint (Using Flask)

We'll use Flask to create a lightweight endpoint that returns the total number of articles from the database. First, install Flask if you haven't already:

pip install flask

Here's how to integrate this into your project (adjust for your specific database setup):

# scrap.py (or a separate app.py file)
from bs4 import BeautifulSoup
import requests
from flask import Flask, jsonify
import sqlite3  # Replace with your DB library (e.g., SQLAlchemy)

app = Flask(__name__)

def get_total_articles():
    """Fetch the total number of articles from the database"""
    conn = sqlite3.connect('articles.db')
    cursor = conn.cursor()
    cursor.execute("SELECT COUNT(*) FROM articles")
    count = cursor.fetchone()[0]
    conn.close()
    return count

# API endpoint to return the count as JSON
@app.route('/api/article-count')
def article_count():
    return jsonify({"total": get_total_articles()})

# Your existing scraping function (updated to save to DB)
def scrape_and_save_article():
    # Example scraping logic (replace with your actual target site)
    url = "https://example.com/new-article"
    response = requests.get(url)
    soup = BeautifulSoup(response.text, 'html.parser')
    
    # Extract article data
    title = soup.find('h1').text.strip()
    content = soup.find('div', class_='article-body').text.strip()
    
    # Save to database
    conn = sqlite3.connect('articles.db')
    cursor = conn.cursor()
    cursor.execute("INSERT INTO articles (title, content) VALUES (?, ?)", (title, content))
    conn.commit()
    conn.close()

if __name__ == "__main__":
    # Run the Flask app (use a WSGI server like Gunicorn for production)
    app.run(debug=True)

2. Frontend: Real-Time Count Update

Now, update your index.html to periodically fetch the count from the API and display it to users. We'll use vanilla JavaScript for this—no extra libraries needed:

<!-- index.html -->
<!DOCTYPE html>
<html lang="en">
<head>
    <meta charset="UTF-8">
    <title>Collected Articles Dashboard</title>
</head>
<body>
    <div class="dashboard">
        <h2>Total Articles Collected: <span id="article-count">0</span></h2>
        <!-- Add other dashboard elements here -->
    </div>

    <script>
        // Function to fetch and update the count
        function updateCount() {
            fetch('/api/article-count')
                .then(response => response.json())
                .then(data => {
                    document.getElementById('article-count').textContent = data.total;
                })
                .catch(error => console.log("Error fetching count:", error));
        }

        // Update count immediately when the page loads
        updateCount();

        // Refresh the count every 3 seconds (adjust this interval as needed)
        setInterval(updateCount, 3000);
    </script>
</body>
</html>

Key Notes to Keep in Mind

  • Database Best Practices: Always close database connections after use to avoid leaks. For larger projects, consider using connection pools.
  • Real-Time Alternatives: If you need instant updates (instead of periodic checks), you could use WebSockets (e.g., with Flask-SocketIO), but the periodic fetch method is simpler for most use cases.
  • Production Setup: Don't use Flask's debug mode in production. Use a proper WSGI server like Gunicorn or uWSGI.

内容的提问来源于stack exchange,提问作者Lotem Nadir

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 07:42:23