如何解决BeautifulSoup无法抓取意铁网站票价数据的问题
The problem you’re hitting is super common with modern websites: the fare information doesn’t come in the initial HTML that requests.get() fetches. Instead, it loads dynamically using JavaScript after the page skeleton is rendered. Since requests doesn’t execute JavaScript, you’re only getting the empty shell of the page—not the actual fare data you need.
Here are two reliable fixes to get your scraper working:
Solution 1: Use Selenium to Simulate a Browser (Easiest for Beginners)
Selenium automates a real web browser (like Chrome or Firefox), which runs all the page’s JavaScript and loads the full dynamic content. Here’s how to adapt your code:
First, install the required packages:
pip install selenium pandas beautifulsoup4
You’ll also need to download ChromeDriver (or Firefox GeckoDriver) and add it to your system PATH.
Then use this code to perform the search and extract fares:
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from bs4 import BeautifulSoup import pandas as pd # Initialize the browser driver = webdriver.Chrome() driver.get("https://www.lefrecce.it/B2CWeb/search.do") try: # Wait for the search form to load, then fill in your details origin_input = WebDriverWait(driver, 10).until( EC.presence_of_element_located((By.ID, "origin")) ) origin_input.send_keys("Torino Porta Nuova") # Your origin station dest_input = driver.find_element(By.ID, "destination") dest_input.send_keys("Milano Centrale") # Your destination station # Set departure date (format: DD/MM/YYYY) date_input = driver.find_element(By.ID, "departureDate") date_input.clear() date_input.send_keys("20/05/2024") # Click the search button search_btn = driver.find_element(By.ID, "searchButton") search_btn.click() # Wait for fares to load on the results page WebDriverWait(driver, 15).until( EC.presence_of_element_located((By.CLASS_NAME, "fare-container")) ) # Grab the fully loaded page source page_source = driver.page_source soup = BeautifulSoup(page_source, "html.parser") # Extract fare data (adjust selectors to match the current page structure) fares = [] for train_row in soup.find_all("div", class_="train-row"): train_details = { "train_number": train_row.find("span", class_="train-number").text.strip(), "departure_time": train_row.find("span", class_="departure-time").text.strip(), "arrival_time": train_row.find("span", class_="arrival-time").text.strip(), "price": train_row.find("span", class_="fare-price").text.strip() } fares.append(train_details) # Convert to a pandas DataFrame for easy handling df = pd.DataFrame(fares) print(df) finally: # Make sure to close the browser when done driver.quit()
Pro Tip: The class names (like train-row or fare-price) might change over time. Use your browser’s dev tools (right-click > Inspect) to check the current HTML structure and update the selectors if needed.
Solution 2: Replicate the API Request (Faster & More Efficient)
Most modern sites load dynamic data via hidden API calls. You can find this using your browser’s Network Tab:
- Go to the search page, fill in your trip details, and click search.
- Open Dev Tools (F12) > Network tab > Filter by "XHR" or "Fetch".
- Look for requests that return fare data (usually in JSON format).
Once you find the API endpoint, you can replicate the request directly with requests:
import requests import pandas as pd # Replace with the actual API endpoint and headers you found in Dev Tools api_url = "https://www.lefrecce.it/B2CWeb/api/v1/search/fares" headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/124.0.0.0 Safari/537.36", "Referer": "https://www.lefrecce.it/B2CWeb/search.do", "X-Requested-With": "XMLHttpRequest" } # Replace with your trip parameters (match what the API expects) payload = { "origin": "TORINO PORTA NUOVA", "destination": "MILANO CENTRALE", "departureDate": "2024-05-20", "adults": 1, # Add any other required parameters from the API request } response = requests.post(api_url, headers=headers, json=payload) data = response.json() # Parse the JSON data into a DataFrame df = pd.DataFrame(data["fares"]) print(df)
This method is faster than Selenium because it skips loading the entire browser, but you’ll need to handle any session tokens or authentication the API requires (check the request headers in Dev Tools for this).
Quick Security Note
Avoid using verify=False in requests.get() unless you have no other option—it disables SSL certificate verification, which exposes you to security risks. If you’re getting SSL errors, make sure your system has the latest root certificates installed.
内容的提问来源于stack exchange,提问作者Shold

