You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

无法获取网站完整HTML响应代码及定时检测预约开放爬取问题求助

Solution for Scraping Berlin Appointment Site & Alerting on Open Slots

Hey there! I’ve run into similar headaches with German municipal service sites before—they’re often tricky because they rely on JavaScript to render content or have basic anti-scraping checks that plain requests can’t bypass. Let’s break down how to fix your problem and build a working monitor.

Why Your requests Code Isn’t Working

The site you’re targeting likely loads the appointment schedule dynamically with JavaScript, not directly in the initial HTML response. requests only fetches raw static HTML, so it misses the content that gets rendered by the browser after the page loads. Additionally, your hardcoded sessionid cookie is probably invalid—sites generate session cookies dynamically when you first visit the homepage, so using a static one won’t work long-term.

Step 1: Use a Headless Browser to Render Full Content

Instead of requests, use a tool like Playwright (my go-to for modern JS-heavy sites) or Selenium. Playwright is lighter, faster, and easier to set up for headless browsing.

First, install Playwright and its browser binaries:

pip install playwright
playwright install chromium

Here’s a sample script to fetch the fully rendered HTML:

from playwright.sync_api import sync_playwright

def get_appointment_page_html():
    with sync_playwright() as p:
        # Launch headless Chromium (set headless=False to see the browser window for testing)
        browser = p.chromium.launch(headless=True)
        context = browser.new_context(
            user_agent='Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/90.0.4430.212 Safari/537.36'
        )
        page = context.new_page()
        
        # First, visit the main appointment page to establish a valid session
        page.goto('https://service.berlin.de/terminvereinbarung/')
        # Then navigate to the day view
        page.goto('https://service.berlin.de/terminvereinbarung/termin/day/')
        
        # Wait for the schedule element to load (adjust the selector to match the actual class/id from the site)
        page.wait_for_selector('.calendar-table')  # Replace with the real selector for the schedule grid
        
        # Get the full rendered HTML
        html_content = page.content()
        browser.close()
        return html_content

Step 2: Check for Open Slots

Once you have the full HTML, parse it with BeautifulSoup to detect available appointments:

from bs4 import BeautifulSoup

def has_open_slots(html_content):
    soup = BeautifulSoup(html_content, 'html.parser')
    # Look for elements that signal open slots (e.g., buttons with "Termin buchen" or highlighted available days)
    # Example selector—inspect the page to find the exact class for available slots:
    open_slot_elements = soup.select('.calendar-day--available')
    return len(open_slot_elements) > 0

Step 3: Set Up Regular Monitoring

Use the schedule library to run your check at intervals (e.g., every 5 minutes). Install it first:

pip install schedule

Add scheduling logic to keep the monitor running:

import schedule
import time
import random

def run_monitor():
    print("Checking for open slots...")
    try:
        html = get_appointment_page_html()
        if has_open_slots(html):
            send_alert()
            print("ALERT: Open slots found!")
        else:
            print("No open slots available.")
        # Add a random delay to mimic human behavior
        time.sleep(random.randint(10, 30))
    except Exception as e:
        print(f"Error during check: {str(e)}")

# Run the check every 5 minutes (adjust as needed—don't check too frequently to avoid being blocked)
schedule.every(5).minutes.do(run_monitor)

# Keep the script running indefinitely
while True:
    schedule.run_pending()
    time.sleep(1)

Step 4: Add Alerting

Choose an alert method that works for you. Here’s how to send an email alert using smtplib (use an app-specific password for Gmail accounts):

import smtplib
from email.mime.text import MIMEText

def send_alert():
    sender_email = 'your-alert-email@gmail.com'
    receiver_email = 'kcanut@gmail.com'
    password = 'your-app-specific-password-here'
    
    msg = MIMEText("Open appointment slots are available! Check here: https://service.berlin.de/terminvereinbarung/termin/day/")
    msg['Subject'] = 'Berlin Appointment Alert: Slots Available!'
    msg['From'] = sender_email
    msg['To'] = receiver_email
    
    with smtplib.SMTP_SSL('smtp.gmail.com', 465) as server:
        server.login(sender_email, password)
        server.sendmail(sender_email, receiver_email, msg.as_string())

Pro Tips to Avoid Getting Blocked

  • Rotate user agents: Occasionally switch the user agent string in Playwright to avoid being flagged as a bot.
  • Don’t overcheck: Stick to 5-10 minute intervals—frequent requests will get your IP blocked.
  • Use a proxy (optional): If you run into IP blocks, use a residential proxy to mask your address.

内容的提问来源于stack exchange,提问作者Koray Can Canut

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.29 17:37:31