You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将统计律师案件数的Python爬虫改造为轻量Web应用?

轻量Web应用实现方案:将KATscrape脚本转为在线工具

Hey there! Since you find Flask and Django too heavy for this specific use case, Bottle is a perfect fit—it's a micro-framework that lives in a single file, has zero extra core dependencies beyond Python itself, and gets the job done with minimal code. Let's walk through adapting your KATscrape script into a fully functional web app:


Step 1: Install Required Packages

First, make sure you have Bottle and your existing scraping dependencies installed:

pip install bottle requests beautifulsoup4 lxml

Step 2: Full Web App Code

Save this single script as kat_scrape_app.py—it combines your scraping logic with a clean web interface:

from bottle import Bottle, request, template
from bs4 import BeautifulSoup
import requests
import re

app = Bottle()

# Your existing court and judge configurations
COURT_NUMBERS = ["1", "7", "8", "4"]
JUDGES_LIST = ["Mr. Justice T.R. Ramachandran Nair", "Mr. V. Somasundaran", "Mr. V.Rajendran", "Mr. Rajesh Dewan", "Mr. Benny Gervacis"]

def scrape_kat_data(date, adv_name):
    """Reusable function to run scraping logic and return structured results"""
    results = []
    # Parse date into day/month/year (expects DD/MM/YYYY format)
    try:
        day, month, year = date.split('/')
    except ValueError:
        return ["Error: Please enter date in DD/MM/YYYY format!"]
    
    base_url = "http://keralaadministrativetribunal.gov.in/ciskat/pages/cause_list_home.php?type=search&dte=%s/%s/%s&court=%s"
    
    for court in COURT_NUMBERS:
        cl_current = base_url % (day, month, year, court)
        try:
            the_page = requests.get(cl_current)
            the_page.raise_for_status()  # Catch HTTP errors
            soup = BeautifulSoup(the_page.content, "lxml")
            da_stuff = str(soup)
            
            # Get sitting judges for the court
            sitting_judges = [judge for judge in JUDGES_LIST if judge in da_stuff]
            
            # Count advocate's cases using regex (word boundary matching)
            regex_pattern = r'\b%s\b' % re.escape(adv_name)
            case_count = sum(1 for _ in re.finditer(regex_pattern, da_stuff))
            
            # Add formatted results
            results.append({
                "court_num": court,
                "presiding": ", ".join(sitting_judges) if sitting_judges else "No judges listed",
                "case_count": case_count,
                "adv_name": adv_name,
                "date": date
            })
        except Exception as e:
            results.append({"court_num": court, "error": f"Failed to scrape: {str(e)}"})
    
    return results

@app.route('/')
def show_search_form():
    """Render the input form for date and advocate name"""
    return template("""
        <!DOCTYPE html>
        <html>
        <head>
            <title>KAT Cause List Search</title>
            <style>
                body { font-family: Arial, sans-serif; max-width: 800px; margin: 2rem auto; padding: 0 1rem; }
                .form-group { margin-bottom: 1.5rem; }
                label { display: block; margin-bottom: 0.5rem; font-weight: bold; }
                input { padding: 0.6rem; width: 100%; max-width: 350px; border: 1px solid #ddd; border-radius: 4px; }
                button { padding: 0.6rem 2rem; background: #2563eb; color: white; border: none; border-radius: 4px; cursor: pointer; }
                button:hover { background: #1d4ed8; }
            </style>
        </head>
        <body>
            <h1>KAT Cause List Search Tool</h1>
            <form action="/search" method="POST">
                <div class="form-group">
                    <label for="date">Date (DD/MM/YYYY):</label>
                    <input type="text" id="date" name="date" required placeholder="Example: 15/03/2024">
                </div>
                <div class="form-group">
                    <label for="adv_name">Advocate Name:</label>
                    <input type="text" id="adv_name" name="adv_name" required placeholder="Example: John Doe">
                </div>
                <button type="submit">Search Cases</button>
            </form>
        </body>
        </html>
    """)

@app.route('/search', method='POST')
def display_results():
    """Handle form submission and render scraping results"""
    date = request.forms.get('date').strip()
    adv_name = request.forms.get('adv_name').strip()
    
    results = scrape_kat_data(date, adv_name)
    
    return template("""
        <!DOCTYPE html>
        <html>
        <head>
            <title>KAT Search Results</title>
            <style>
                body { font-family: Arial, sans-serif; max-width: 800px; margin: 2rem auto; padding: 0 1rem; }
                .result-section { margin-bottom: 1.5rem; padding: 1rem; border: 1px solid #ddd; border-radius: 4px; }
                .error { color: #dc2626; }
                .back-link { margin-top: 2rem; display: inline-block; color: #2563eb; text-decoration: none; }
                .back-link:hover { text-decoration: underline; }
            </style>
        </head>
        <body>
            <h1>Results for {{adv_name}} on {{date}}</h1>
            <a href="/" class="back-link">← Go Back to Search</a>
            <div class="results-container">
                % for result in results:
                    <div class="result-section">
                        <h3>Court No. {{result.court_num}}</h3>
                        % if 'error' in result:
                            <p class="error">{{result.error}}</p>
                        % else:
                            <p><strong>Presiding Judges:</strong> {{result.presiding}}</p>
                            <p><strong>Total Matters Assigned:</strong> {{result.case_count}}</p>
                        % end
                    </div>
                % end
            </div>
        </body>
        </html>
    """, results=results, adv_name=adv_name, date=date)

if __name__ == '__main__':
    app.run(host='localhost', port=8080, debug=True)

Step 3: How It Works

  1. Input Form: Visit http://localhost:8080 after running the script to access a simple form where users enter the date and advocate's name.
  2. Scraping Logic: The scrape_kat_data function is a cleaned-up version of your original script—it takes user inputs, runs the scraping for each court, and returns structured results instead of printing directly.
  3. Results Page: When the form is submitted, the app processes the inputs, runs the scraper, and displays formatted, easy-to-read results on a new page.

Why This Is Lightweight

  • Single File: The entire app lives in one script—no complex project folders or configuration files.
  • Minimal Overhead: Bottle has almost zero runtime overhead, so the app launches instantly.
  • Familiar Logic: Your original scraping code is preserved and reused, so you don't have to rewrite core functionality.

Alternative Lightweight Options

If you want to explore other tools:

  • FastAPI: Slightly more feature-rich but still lightweight, with automatic API documentation (great if you ever want to add an API endpoint).
  • CherryPy: A micro-framework with a built-in server and straightforward routing, ideal for small web tools.

Just run the script, open your browser to http://localhost:8080, and you're ready to use your online KAT cause list search tool!

内容的提问来源于stack exchange,提问作者Basil Ajith

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.06 22:57:45