Python新手求助:荷兰透明度基准网站动态数据抓取问题
Hey there! As a Python newbie hitting this dynamic content roadblock, I totally get how frustrating it can be—regular requests-based scraping just can’t grab data that loads after the initial page renders. Let’s break down two reliable approaches to get those company details you need:
1. Use Browser Automation Tools (Selenium or Playwright)
These tools simulate a real browser, so they’ll wait for dynamic content to load just like your regular browser does. Here’s a step-by-step example with Selenium, which is pretty beginner-friendly:
Step 1: Install Required Tools
First, install the Selenium package and a web driver matching your browser version (e.g., ChromeDriver for Chrome):
pip install selenium
Step 2: Write the Scraping Code
This script will launch Chrome, load the page, wait for company data to appear, then extract the fully rendered content:
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from bs4 import BeautifulSoup # Initialize the Chrome browser driver = webdriver.Chrome() # Navigate to the target page (replace with your actual URL) driver.get("YOUR_TARGET_PAGE_URL") try: # Wait up to 10 seconds for the company list to load (adjust selector to match the site) WebDriverWait(driver, 10).until( EC.presence_of_element_located((By.CLASS_NAME, "company-item")) ) # Grab the fully rendered HTML rendered_html = driver.page_source # Parse with BeautifulSoup to extract details soup = BeautifulSoup(rendered_html, "html.parser") companies = soup.find_all(class_="company-item") for company in companies: # Extract specific info (adjust based on the site's structure) name = company.find("h3").text.strip() score = company.find(class_="transparency-score").text.strip() print(f"Company: {name} | Transparency Score: {score}") finally: # Make sure to close the browser when done driver.quit()
Note: Replace "YOUR_TARGET_PAGE_URL" and the element selectors (like "company-item") with values you find via browser DevTools.
2. Find and Call the Underlying API
Most dynamic sites load data via hidden API requests (XHR/Fetch). This method is faster than browser automation because you don’t need to run a full browser. Here’s how to do it:
Step 1: Locate the API Request
- Open your browser’s DevTools (F12) and switch to the Network tab.
- Refresh the page and filter for "XHR" or "Fetch" requests.
- Click through these requests to find the one that returns company data (check the "Response" tab to confirm).
- Copy the request URL, headers, and any required parameters (like page numbers or filters).
Step 2: Fetch Data with requests
Use the requests library to call the API directly:
import requests # Replace with the actual API endpoint you found api_url = "https://example-api-endpoint.com/transparency/companies" # Add necessary headers (often a User-Agent to mimic a real browser) headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36" } # Add any required query parameters (e.g., page number, sector filter) params = { "page": 1, "sector": "all" } response = requests.get(api_url, headers=headers, params=params) if response.status_code == 200: # Parse the JSON response company_data = response.json() for company in company_data["results"]: print(f"Name: {company['name']} | Score: {company['score']}") else: print(f"Request failed with status code: {response.status_code}")
Quick Tips for Newbies
- Always check the site’s
robots.txtfile (add/robots.txtto the site URL) to ensure scraping is allowed. - Add small delays between requests (
import time; time.sleep(2)) to avoid overwhelming the server and getting blocked. - For browser automation, use headless mode (e.g.,
options.add_argument("--headless=new")for Chrome) to run the browser in the background.
内容的提问来源于stack exchange,提问作者I_love_Norway

