求助:如何获取可抓取不同赔率值的有效XPath?脚本执行异常
Let's break down why your XPath works in isolation but returns duplicate values when integrated into your script, and walk through actionable fixes:
1. Fix Contextual Selection in Your Script
The most likely culprit is that your script is reusing the same parent element context for every selection, instead of iterating through each unique matching node from your target XPath.
For example, if you're doing something like this (which causes duplicates):
# Wrong approach - reusing a single parent context parent = driver.find_element(By.XPATH, "//div[contains(@class, 'gl-Market_HasLabels')]") odds = parent.find_elements(By.XPATH, "//div[contains(@class, 'gl-ParticipantOddsOnly')]//span[@class='gl-ParticipantOddsOnly_Odds']")
Instead, first collect all the target odds containers, then iterate through each one to extract values using a relative XPath (starting with . to stay within the current container):
# Correct approach - iterate through each matched node odds_containers = driver.find_elements(By.XPATH, "//div[contains(@class, 'gl-Market_HasLabels')]/following-sibling::div[contains(@class, 'gl-Market_PWidth-12-3333')][1]//div[contains(@class, 'gl-ParticipantOddsOnly')]") all_odds = [] for container in odds_containers: # Use .// to target elements inside the current container only odds_value = container.find_element(By.XPATH, ".//span[@class='gl-ParticipantOddsOnly_Odds']").text all_odds.append(odds_value) print(all_odds)
2. Refine Your XPath for Precision
Your original XPath is a bit broad. Let's tighten it to directly target the span holding the odds (since your sample HTML shows it has an exact class gl-ParticipantOddsOnly_Odds):
//div[contains(@class, 'gl-Market_HasLabels')]/following-sibling::div[contains(@class, 'gl-Market_PWidth-12-3333')][1]//div[contains(@class, 'gl-ParticipantOddsOnly')]/span[@class='gl-ParticipantOddsOnly_Odds']
With this XPath, you can collect all odds elements directly and extract their text without extra steps:
odds_elements = driver.find_elements(By.XPATH, "//div[contains(@class, 'gl-Market_HasLabels')]/following-sibling::div[contains(@class, 'gl-Market_PWidth-12-3333')][1]//div[contains(@class, 'gl-ParticipantOddsOnly')]/span[@class='gl-ParticipantOddsOnly_Odds']") all_odds = [elem.text for elem in odds_elements]
3. Ensure Group and Odds Length Consistency
You noted that groups and xp_ba3 need matching lengths. To avoid mismatches, collect both sets of elements from the same parent market section:
# Target the specific market section first market_section = driver.find_element(By.XPATH, "//div[contains(@class, 'gl-Market_HasLabels')]/following-sibling::div[contains(@class, 'gl-Market_PWidth-12-3333')][1]") # Collect groups and odds from this section groups = market_section.find_elements(By.XPATH, ".//div[contains(@class, 'gl-Participant')]") odds = market_section.find_elements(By.XPATH, ".//span[@class='gl-ParticipantOddsOnly_Odds']") if len(groups) == len(odds): for group, odd in zip(groups, odds): print(f"{group.text}: {odd.text}") else: print(f"Mismatch alert: {len(groups)} groups vs {len(odds)} odds")
4. Account for Dynamic Content Loading
If the page loads odds via JavaScript, your script might extract values before the page fully updates. Add explicit waits to ensure elements are ready:
from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC # Wait up to 10 seconds for at least one odds element to load wait = WebDriverWait(driver, 10) wait.until(EC.presence_of_element_located((By.XPATH, "//span[@class='gl-ParticipantOddsOnly_Odds']"))) # Now scrape the odds odds_elements = driver.find_elements(By.XPATH, "//div[contains(@class, 'gl-Market_HasLabels')]/following-sibling::div[contains(@class, 'gl-Market_PWidth-12-3333')][1]//span[@class='gl-ParticipantOddsOnly_Odds']") all_odds = [elem.text for elem in odds_elements]
5. Debug with HTML Inspections
Add debug prints to verify what your script is actually selecting. This helps catch hidden duplicate elements or incorrect context:
odds_containers = driver.find_elements(By.XPATH, "//div[contains(@class, 'gl-ParticipantOddsOnly')]") print(f"Found {len(odds_containers)} odds containers") for i, container in enumerate(odds_containers): print(f"Container {i} HTML: {container.get_attribute('outerHTML')}")
内容的提问来源于stack exchange,提问作者user9155788

