基于Selenium优化足球赛事数据爬取的XPath问题咨询
Fixing Flashscore Scraping: Half-Time Results, Goal Times & Half-Time Stats
Let’s get your scraping code working reliably with updated selectors that won’t break easily when Flashscore tweaks their page structure. I’ve revised your code to properly capture all the data you need, plus cleaned up some logic gaps:
Revised Python Code
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from selenium.common.exceptions import NoSuchElementException import pandas as pd stats = [] driver = webdriver.Chrome() # Swap for your preferred driver (Firefox, Edge, etc.) try: # Load match summary page driver.get('https://www.flashscore.co.uk/match/dxm5iiWh/#match-summary') # Get match date/time match_datetime = WebDriverWait(driver, 10).until( EC.visibility_of_element_located((By.CLASS_NAME, "duelParticipant__startTime")) ).text # Get home and away team names home_team = WebDriverWait(driver, 10).until( EC.visibility_of_element_located((By.CLASS_NAME, "participantHome__name")) ).text away_team = WebDriverWait(driver, 10).until( EC.visibility_of_element_located((By.CLASS_NAME, "participantAway__name")) ).text # Collect all goal times, split into first/second half all_goals = [] incident_rows = WebDriverWait(driver, 10).until( EC.presence_of_all_elements_located((By.CLASS_NAME, "detailMS__incidentRow")) ) for row in incident_rows: # Check if the row is a goal (has the soccer ball icon) if row.find_elements(By.CLASS_NAME, "icon--soccer-ball"): goal_details = row.text.strip() all_goals.append(goal_details) # Split goals into first/second half for cleaner data first_half_goals = [g for g in all_goals if "'" in g and int(g.split("'")[0]) <= 45] second_half_goals = [g for g in all_goals if "'" in g and int(g.split("'")[0]) > 45] goals_summary = f"First Half: {', '.join(first_half_goals) if first_half_goals else 'No goals'} | Second Half: {', '.join(second_half_goals) if second_half_goals else 'No goals'}" # Get extra match info (referee, stadium, etc.) try: extra_info = WebDriverWait(driver, 5).until( EC.visibility_of_element_located((By.CLASS_NAME, "detailMS__extraInfo")) ).text except NoSuchElementException: extra_info = "No extra info available" # Navigate to half-time statistics tab try: stats_tab = WebDriverWait(driver, 10).until( EC.element_to_be_clickable((By.ID, "a-match-statistics")) ) stats_tab.click() # Wait for half-time stats to load WebDriverWait(driver, 10).until( EC.visibility_of_element_located((By.ID, "tab-statistics-0-statistic")) ) # Extract shots (half-time) home_shots = WebDriverWait(driver, 10).until( EC.visibility_of_element_located((By.XPATH, "//div[@id='tab-statistics-0-statistic']//div[contains(text(), 'Shots')]/following-sibling::div/div[1]")) ).text away_shots = WebDriverWait(driver, 10).until( EC.visibility_of_element_located((By.XPATH, "//div[@id='tab-statistics-0-statistic']//div[contains(text(), 'Shots')]/following-sibling::div/div[3]")) ).text # Extract shots on target (half-time) home_shot_on_target = WebDriverWait(driver, 10).until( EC.visibility_of_element_located((By.XPATH, "//div[@id='tab-statistics-0-statistic']//div[contains(text(), 'Shots on target')]/following-sibling::div/div[1]")) ).text away_shot_on_target = WebDriverWait(driver, 10).until( EC.visibility_of_element_located((By.XPATH, "//div[@id='tab-statistics-0-statistic']//div[contains(text(), 'Shots on target')]/following-sibling::div/div[3]")) ).text except NoSuchElementException: home_shots = "No stats available" away_shots = "No stats available" home_shot_on_target = "No stats available" away_shot_on_target = "No stats available" # Add all collected data to the stats list stats.append({ "match_datetime": match_datetime, "home_team": home_team, "away_team": away_team, "goals_summary": goals_summary, "extra_info": extra_info, "home_half_shots": home_shots, "away_half_shots": away_shots, "home_half_shot_on_target": home_shot_on_target, "away_half_shot_on_target": away_shot_on_target }) finally: driver.quit() # Ensure the browser closes even if an error occurs # Convert to DataFrame and save to CSV df = pd.DataFrame(stats) df.to_csv('flashscore_match_stats.csv', index=False, encoding='utf-8') print("Data saved to flashscore_match_stats.csv successfully!")
Key Fixes & Improvements
- Stable Selectors: Replaced fragile absolute XPaths with class-based selectors and relative XPaths (using
contains()for stat rows) that’re far less likely to break when Flashscore updates their UI. - Complete Goal Capture: Instead of only grabbing the last goal, we collect all goals and split them into first/second half for more organized data.
- Robust Error Handling: Added a
try/finallyblock to guarantee the browser driver closes properly, even if a scrape fails mid-process. - Cleaner Data Structure: Used a dictionary to organize collected data, making the final CSV columns more readable.
- Explicit Waits: Increased wait times slightly (to 10 seconds) and used
presence_of_all_elements_locatedfor incident rows to avoid missing data due to slow page loads.
To expand to more half-time stats, just replicate the pattern used for shots/shots on target. For example, to add possession:
home_possession = WebDriverWait(driver, 10).until( EC.visibility_of_element_located((By.XPATH, "//div[@id='tab-statistics-0-statistic']//div[contains(text(), 'Possession')]/following-sibling::div/div[1]")) ).text
内容的提问来源于stack exchange,提问作者vriones11
相关产品推荐
相关产品推荐

