You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python新手求助:荷兰透明度基准网站动态数据抓取问题

Solutions for Scraping Dynamic Data from Dutch Transparency Benchmark

Hey there! As a Python newbie hitting this dynamic content roadblock, I totally get how frustrating it can be—regular requests-based scraping just can’t grab data that loads after the initial page renders. Let’s break down two reliable approaches to get those company details you need:

1. Use Browser Automation Tools (Selenium or Playwright)

These tools simulate a real browser, so they’ll wait for dynamic content to load just like your regular browser does. Here’s a step-by-step example with Selenium, which is pretty beginner-friendly:

Step 1: Install Required Tools

First, install the Selenium package and a web driver matching your browser version (e.g., ChromeDriver for Chrome):

pip install selenium

Step 2: Write the Scraping Code

This script will launch Chrome, load the page, wait for company data to appear, then extract the fully rendered content:

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from bs4 import BeautifulSoup

# Initialize the Chrome browser
driver = webdriver.Chrome()

# Navigate to the target page (replace with your actual URL)
driver.get("YOUR_TARGET_PAGE_URL")

try:
    # Wait up to 10 seconds for the company list to load (adjust selector to match the site)
    WebDriverWait(driver, 10).until(
        EC.presence_of_element_located((By.CLASS_NAME, "company-item"))
    )
    # Grab the fully rendered HTML
    rendered_html = driver.page_source
    # Parse with BeautifulSoup to extract details
    soup = BeautifulSoup(rendered_html, "html.parser")
    companies = soup.find_all(class_="company-item")
    
    for company in companies:
        # Extract specific info (adjust based on the site's structure)
        name = company.find("h3").text.strip()
        score = company.find(class_="transparency-score").text.strip()
        print(f"Company: {name} | Transparency Score: {score}")
finally:
    # Make sure to close the browser when done
    driver.quit()

Note: Replace "YOUR_TARGET_PAGE_URL" and the element selectors (like "company-item") with values you find via browser DevTools.

2. Find and Call the Underlying API

Most dynamic sites load data via hidden API requests (XHR/Fetch). This method is faster than browser automation because you don’t need to run a full browser. Here’s how to do it:

Step 1: Locate the API Request

  1. Open your browser’s DevTools (F12) and switch to the Network tab.
  2. Refresh the page and filter for "XHR" or "Fetch" requests.
  3. Click through these requests to find the one that returns company data (check the "Response" tab to confirm).
  4. Copy the request URL, headers, and any required parameters (like page numbers or filters).

Step 2: Fetch Data with requests

Use the requests library to call the API directly:

import requests

# Replace with the actual API endpoint you found
api_url = "https://example-api-endpoint.com/transparency/companies"

# Add necessary headers (often a User-Agent to mimic a real browser)
headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36"
}

# Add any required query parameters (e.g., page number, sector filter)
params = {
    "page": 1,
    "sector": "all"
}

response = requests.get(api_url, headers=headers, params=params)

if response.status_code == 200:
    # Parse the JSON response
    company_data = response.json()
    for company in company_data["results"]:
        print(f"Name: {company['name']} | Score: {company['score']}")
else:
    print(f"Request failed with status code: {response.status_code}")

Quick Tips for Newbies

  • Always check the site’s robots.txt file (add /robots.txt to the site URL) to ensure scraping is allowed.
  • Add small delays between requests (import time; time.sleep(2)) to avoid overwhelming the server and getting blocked.
  • For browser automation, use headless mode (e.g., options.add_argument("--headless=new") for Chrome) to run the browser in the background.

内容的提问来源于stack exchange,提问作者I_love_Norway

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 09:35:04