无法获取网站完整HTML响应代码及定时检测预约开放爬取问题求助
Hey there! I’ve run into similar headaches with German municipal service sites before—they’re often tricky because they rely on JavaScript to render content or have basic anti-scraping checks that plain requests can’t bypass. Let’s break down how to fix your problem and build a working monitor.
Why Your requests Code Isn’t Working
The site you’re targeting likely loads the appointment schedule dynamically with JavaScript, not directly in the initial HTML response. requests only fetches raw static HTML, so it misses the content that gets rendered by the browser after the page loads. Additionally, your hardcoded sessionid cookie is probably invalid—sites generate session cookies dynamically when you first visit the homepage, so using a static one won’t work long-term.
Step 1: Use a Headless Browser to Render Full Content
Instead of requests, use a tool like Playwright (my go-to for modern JS-heavy sites) or Selenium. Playwright is lighter, faster, and easier to set up for headless browsing.
First, install Playwright and its browser binaries:
pip install playwright playwright install chromium
Here’s a sample script to fetch the fully rendered HTML:
from playwright.sync_api import sync_playwright def get_appointment_page_html(): with sync_playwright() as p: # Launch headless Chromium (set headless=False to see the browser window for testing) browser = p.chromium.launch(headless=True) context = browser.new_context( user_agent='Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/90.0.4430.212 Safari/537.36' ) page = context.new_page() # First, visit the main appointment page to establish a valid session page.goto('https://service.berlin.de/terminvereinbarung/') # Then navigate to the day view page.goto('https://service.berlin.de/terminvereinbarung/termin/day/') # Wait for the schedule element to load (adjust the selector to match the actual class/id from the site) page.wait_for_selector('.calendar-table') # Replace with the real selector for the schedule grid # Get the full rendered HTML html_content = page.content() browser.close() return html_content
Step 2: Check for Open Slots
Once you have the full HTML, parse it with BeautifulSoup to detect available appointments:
from bs4 import BeautifulSoup def has_open_slots(html_content): soup = BeautifulSoup(html_content, 'html.parser') # Look for elements that signal open slots (e.g., buttons with "Termin buchen" or highlighted available days) # Example selector—inspect the page to find the exact class for available slots: open_slot_elements = soup.select('.calendar-day--available') return len(open_slot_elements) > 0
Step 3: Set Up Regular Monitoring
Use the schedule library to run your check at intervals (e.g., every 5 minutes). Install it first:
pip install schedule
Add scheduling logic to keep the monitor running:
import schedule import time import random def run_monitor(): print("Checking for open slots...") try: html = get_appointment_page_html() if has_open_slots(html): send_alert() print("ALERT: Open slots found!") else: print("No open slots available.") # Add a random delay to mimic human behavior time.sleep(random.randint(10, 30)) except Exception as e: print(f"Error during check: {str(e)}") # Run the check every 5 minutes (adjust as needed—don't check too frequently to avoid being blocked) schedule.every(5).minutes.do(run_monitor) # Keep the script running indefinitely while True: schedule.run_pending() time.sleep(1)
Step 4: Add Alerting
Choose an alert method that works for you. Here’s how to send an email alert using smtplib (use an app-specific password for Gmail accounts):
import smtplib from email.mime.text import MIMEText def send_alert(): sender_email = 'your-alert-email@gmail.com' receiver_email = 'kcanut@gmail.com' password = 'your-app-specific-password-here' msg = MIMEText("Open appointment slots are available! Check here: https://service.berlin.de/terminvereinbarung/termin/day/") msg['Subject'] = 'Berlin Appointment Alert: Slots Available!' msg['From'] = sender_email msg['To'] = receiver_email with smtplib.SMTP_SSL('smtp.gmail.com', 465) as server: server.login(sender_email, password) server.sendmail(sender_email, receiver_email, msg.as_string())
Pro Tips to Avoid Getting Blocked
- Rotate user agents: Occasionally switch the user agent string in Playwright to avoid being flagged as a bot.
- Don’t overcheck: Stick to 5-10 minute intervals—frequent requests will get your IP blocked.
- Use a proxy (optional): If you run into IP blocks, use a residential proxy to mask your address.
内容的提问来源于stack exchange,提问作者Koray Can Canut

