如何将统计律师案件数的Python爬虫改造为轻量Web应用?
Hey there! Since you find Flask and Django too heavy for this specific use case, Bottle is a perfect fit—it's a micro-framework that lives in a single file, has zero extra core dependencies beyond Python itself, and gets the job done with minimal code. Let's walk through adapting your KATscrape script into a fully functional web app:
Step 1: Install Required Packages
First, make sure you have Bottle and your existing scraping dependencies installed:
pip install bottle requests beautifulsoup4 lxml
Step 2: Full Web App Code
Save this single script as kat_scrape_app.py—it combines your scraping logic with a clean web interface:
from bottle import Bottle, request, template from bs4 import BeautifulSoup import requests import re app = Bottle() # Your existing court and judge configurations COURT_NUMBERS = ["1", "7", "8", "4"] JUDGES_LIST = ["Mr. Justice T.R. Ramachandran Nair", "Mr. V. Somasundaran", "Mr. V.Rajendran", "Mr. Rajesh Dewan", "Mr. Benny Gervacis"] def scrape_kat_data(date, adv_name): """Reusable function to run scraping logic and return structured results""" results = [] # Parse date into day/month/year (expects DD/MM/YYYY format) try: day, month, year = date.split('/') except ValueError: return ["Error: Please enter date in DD/MM/YYYY format!"] base_url = "http://keralaadministrativetribunal.gov.in/ciskat/pages/cause_list_home.php?type=search&dte=%s/%s/%s&court=%s" for court in COURT_NUMBERS: cl_current = base_url % (day, month, year, court) try: the_page = requests.get(cl_current) the_page.raise_for_status() # Catch HTTP errors soup = BeautifulSoup(the_page.content, "lxml") da_stuff = str(soup) # Get sitting judges for the court sitting_judges = [judge for judge in JUDGES_LIST if judge in da_stuff] # Count advocate's cases using regex (word boundary matching) regex_pattern = r'\b%s\b' % re.escape(adv_name) case_count = sum(1 for _ in re.finditer(regex_pattern, da_stuff)) # Add formatted results results.append({ "court_num": court, "presiding": ", ".join(sitting_judges) if sitting_judges else "No judges listed", "case_count": case_count, "adv_name": adv_name, "date": date }) except Exception as e: results.append({"court_num": court, "error": f"Failed to scrape: {str(e)}"}) return results @app.route('/') def show_search_form(): """Render the input form for date and advocate name""" return template(""" <!DOCTYPE html> <html> <head> <title>KAT Cause List Search</title> <style> body { font-family: Arial, sans-serif; max-width: 800px; margin: 2rem auto; padding: 0 1rem; } .form-group { margin-bottom: 1.5rem; } label { display: block; margin-bottom: 0.5rem; font-weight: bold; } input { padding: 0.6rem; width: 100%; max-width: 350px; border: 1px solid #ddd; border-radius: 4px; } button { padding: 0.6rem 2rem; background: #2563eb; color: white; border: none; border-radius: 4px; cursor: pointer; } button:hover { background: #1d4ed8; } </style> </head> <body> <h1>KAT Cause List Search Tool</h1> <form action="/search" method="POST"> <div class="form-group"> <label for="date">Date (DD/MM/YYYY):</label> <input type="text" id="date" name="date" required placeholder="Example: 15/03/2024"> </div> <div class="form-group"> <label for="adv_name">Advocate Name:</label> <input type="text" id="adv_name" name="adv_name" required placeholder="Example: John Doe"> </div> <button type="submit">Search Cases</button> </form> </body> </html> """) @app.route('/search', method='POST') def display_results(): """Handle form submission and render scraping results""" date = request.forms.get('date').strip() adv_name = request.forms.get('adv_name').strip() results = scrape_kat_data(date, adv_name) return template(""" <!DOCTYPE html> <html> <head> <title>KAT Search Results</title> <style> body { font-family: Arial, sans-serif; max-width: 800px; margin: 2rem auto; padding: 0 1rem; } .result-section { margin-bottom: 1.5rem; padding: 1rem; border: 1px solid #ddd; border-radius: 4px; } .error { color: #dc2626; } .back-link { margin-top: 2rem; display: inline-block; color: #2563eb; text-decoration: none; } .back-link:hover { text-decoration: underline; } </style> </head> <body> <h1>Results for {{adv_name}} on {{date}}</h1> <a href="/" class="back-link">← Go Back to Search</a> <div class="results-container"> % for result in results: <div class="result-section"> <h3>Court No. {{result.court_num}}</h3> % if 'error' in result: <p class="error">{{result.error}}</p> % else: <p><strong>Presiding Judges:</strong> {{result.presiding}}</p> <p><strong>Total Matters Assigned:</strong> {{result.case_count}}</p> % end </div> % end </div> </body> </html> """, results=results, adv_name=adv_name, date=date) if __name__ == '__main__': app.run(host='localhost', port=8080, debug=True)
Step 3: How It Works
- Input Form: Visit
http://localhost:8080after running the script to access a simple form where users enter the date and advocate's name. - Scraping Logic: The
scrape_kat_datafunction is a cleaned-up version of your original script—it takes user inputs, runs the scraping for each court, and returns structured results instead of printing directly. - Results Page: When the form is submitted, the app processes the inputs, runs the scraper, and displays formatted, easy-to-read results on a new page.
Why This Is Lightweight
- Single File: The entire app lives in one script—no complex project folders or configuration files.
- Minimal Overhead: Bottle has almost zero runtime overhead, so the app launches instantly.
- Familiar Logic: Your original scraping code is preserved and reused, so you don't have to rewrite core functionality.
Alternative Lightweight Options
If you want to explore other tools:
- FastAPI: Slightly more feature-rich but still lightweight, with automatic API documentation (great if you ever want to add an API endpoint).
- CherryPy: A micro-framework with a built-in server and straightforward routing, ideal for small web tools.
Just run the script, open your browser to http://localhost:8080, and you're ready to use your online KAT cause list search tool!
内容的提问来源于stack exchange,提问作者Basil Ajith

