Selenium爬取赛马分段时间数据缺失11条记录的问题排查与解决方案求助
Selenium爬取赛马分段时间数据缺失11条记录的问题排查与解决方案求助
各位大佬好!我刚接触Python 3.13.2,最近在用Selenium配合Firefox爬取赛马赛事页面的「分段时间」数据,目标是切尔滕纳姆某场赛事的详情页。我的需求很明确:收集11匹马的分段时间,每匹马对应8个数据点,总共应该有88条记录,但目前只拿到了77条,整整少了11条。我怀疑是自己的过滤规则把部分马匹的相关数据误筛掉了,实在卡在这里了,想请大家帮忙出出主意!
我主要有几个疑问:
- 怎么调整过滤规则,才能拿到全部88条分段时间?尤其是有些时间只有一位小数的情况
- 我现在是跳过前69个单元格来避开表头,这种做法靠谱吗?有没有更精准的方式定位到马匹数据的起始位置?
- 有没有办法确保11匹马的数据都被完整处理,哪怕部分马匹的分段时间存在缺失?
接下来跟大家说说我目前的操作流程和遇到的问题:
我先通过Selenium加载页面,点击「分段时间」标签,然后反复滚动页面确保所有内容都加载完成,总共抓取到了168个单元格。为了筛选出有效的分段时间,我设置了规则:只保留符合格式的小数(比如61.58或6.5),并且排除10及以下的数值——我以为这些是马匹编号(比如5)。最后处理下来只拿到了9匹马的数据,每匹补全8个分段时间(缺失的用0.00填充)。
预期结果:拿到88条分段时间,11匹马每匹对应8个有效数据,形成完整的数据集。
实际结果:只拿到77个有效单元格,仅处理了9匹马。
我收集到的调试信息:
All time cells found: 168 First 10 raw time cells: ['Pos', 'Silk', 'Horse', '1', '', ...] Filtered time cells found: 77 Accepted samples: ['61.58', '56.55', '30.02', '27.18', '26.97'] Rejected samples: ['5.', '11.', '6.', '4.', '7.', '10.', '2.', '8.', '1.', '3.'] Horses Found in final data: 9
我的代码尝试:
from selenium import webdriver from selenium.webdriver.firefox.service import Service from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC import pandas as pd import time # Firefox service service = Service(executable_path="C:\\Scraping\\geckodriver.exe") driver = webdriver.Firefox(service=service) try: # Racecard URL url = "https://www.attheraces.com/racecard/Cheltenham/11-March-2025/1320" driver.get(url) time.sleep(15) # Extended wait to allow full page load driver.save_screenshot("C:\\Scraping\\initial_load.png") print("Initial load captured") # Debug initial page source with open("C:\\Scraping\\page_source_initial.html", "w", encoding="utf-8") as f: f.write(driver.page_source) print("Initial page source saved") # Accept cookies try: cookie_button = WebDriverWait(driver, 10).until( EC.element_to_be_clickable((By.XPATH, '//button[text()="Accept All"]')) ) cookie_button.click() print("Cookies accepted") time.sleep(5) except Exception as e: print("No cookie prompt found:", e) # Click Sectional Times tab with presence check and forced click print("Attempting to click Sectional Times tab...") try: tab = WebDriverWait(driver, 40).until( EC.presence_of_element_located((By.XPATH, '//a[contains(@class, "tab") and contains(normalize-space(.), "Sectional")]')) ) driver.execute_script("arguments[0].click();", tab) print("Tab clicked") except Exception as e: print(f"Tab not found or click failed: {e}") driver.save_screenshot("C:\\Scraping\\tab_click_error.png") time.sleep(15) # Wait for tab container print("Waiting for tab container...") WebDriverWait(driver, 30).until( EC.presence_of_element_located((By.XPATH, "//div[@id='tab-sectionals-times']")) ) print("Tab container found") # Multiple scroll attempts to load content print("Scrolling to load content...") for _ in range(5): driver.execute_script("window.scrollTo(0, document.body.scrollHeight);") time.sleep(5) driver.execute_script("window.scrollTo(0, 0);") time.sleep(2) driver.execute_script("window.scrollTo(0, document.body.scrollHeight);") time.sleep(5) driver.save_screenshot("C:\\Scraping\\after_scroll.png") print("Scroll completed") # Wait for sectional times with retry and extended timeout print("Waiting for sectional times...") max_attempts = 3 for attempt in range(max_attempts): try: WebDriverWait(driver, 90).until( # Increased to 90 seconds EC.presence_of_all_elements_located((By.XPATH, "//div[@id='tab-sectionals-times']//div[contains(@class, 'card-cell')]")) ) print("Sectional times found") break except Exception as e: print(f"Attempt {attempt + 1} failed: {e}") if attempt < max_attempts - 1: print("Retrying...") time.sleep(10) else: print("Max retries reached, raising error") driver.save_screenshot("C:\\Scraping\\error_screenshot.png") raise driver.save_screenshot("C:\\Scraping\\after_times_wait.png") # Debug page source with open("C:\\Scraping\\page_source_after_scroll.html", "w", encoding="utf-8") as f: f.write(driver.page_source) # Headers headers = ["start-12f", "12f-8f", "8f-6f", "6f-4f", "4f-2f", "2f-1f", "1f-finish", "finish"] all_data = [] # Race info race_name = driver.find_element(By.TAG_NAME, "h1").text.strip() if driver.find_elements(By.TAG_NAME, "h1") else "Unknown Race" distance = driver.find_element(By.CLASS_NAME, "p--large.font-weight--semibold").text.strip() if driver.find_elements(By.CLASS_NAME, "p--large.font-weight--semibold") else "0f" meeting = "cheltenham" # Scrape horse names print("Scraping horse names...") horse_elements = driver.find_elements(By.XPATH, "//div[contains(@class, 'card-entry')]//h2//a[contains(@class, 'horse__link')]") horses = [elem.text.strip() for elem in horse_elements if elem.text.strip() and elem.is_displayed()] print("Horses found:", len(horses)) # Scrape sectional times with corrected alignment print("Scraping sectional times...") time_cells = driver.find_elements(By.XPATH, "//div[@id='tab-sectionals-times']//div[contains(@class, 'card-cell')]") print("All time cells found:", len(time_cells)) print("First 10 raw time cells:", [cell.text.strip() for cell in time_cells[:10]]) # Debug raw data # Filter for time cells, accepting cleanable decimals after skipping headers time_cells_filtered = [] accepted_samples = [] rejected_samples = [] header_count = 69 # Estimated based on raw data pattern for i, cell in enumerate(time_cells[header_count:]): # Skip initial headers spans = cell.find_elements(By.XPATH, ".//span[contains(@class, 'visible')]") for span in spans: text = span.text.strip() cleaned = text.replace(' ', '').replace('x', '') if '.' in text and cleaned.replace('.', '').isdigit(): parts = cleaned.split('.') if len(parts) == 2 and all(part.isdigit() for part in parts) and len(parts[0]) in [1, 2] and len(parts[1]) in [1, 2] and float(cleaned) > 10: time_cells_filtered.append(cell) accepted_samples.append(text) else: rejected_samples.append(text) print("Filtered time cells found:", len(time_cells_filtered)) if accepted_samples: print("Accepted samples:", accepted_samples[:5]) # Show first 5 accepted if rejected_samples: print("Rejected samples:", rejected_samples[:10]) # Show first 10 rejected
备注:内容来源于stack exchange,提问作者Audrey E Unique Expressions Ar
相关产品推荐
相关产品推荐

