You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Selenium爬取赛马分段时间数据缺失11条记录的问题排查与解决方案求助

Selenium爬取赛马分段时间数据缺失11条记录的问题排查与解决方案求助

各位大佬好!我刚接触Python 3.13.2,最近在用Selenium配合Firefox爬取赛马赛事页面的「分段时间」数据,目标是切尔滕纳姆某场赛事的详情页。我的需求很明确:收集11匹马的分段时间,每匹马对应8个数据点,总共应该有88条记录,但目前只拿到了77条,整整少了11条。我怀疑是自己的过滤规则把部分马匹的相关数据误筛掉了,实在卡在这里了,想请大家帮忙出出主意!

我主要有几个疑问:

  • 怎么调整过滤规则,才能拿到全部88条分段时间?尤其是有些时间只有一位小数的情况
  • 我现在是跳过前69个单元格来避开表头,这种做法靠谱吗?有没有更精准的方式定位到马匹数据的起始位置?
  • 有没有办法确保11匹马的数据都被完整处理,哪怕部分马匹的分段时间存在缺失?

接下来跟大家说说我目前的操作流程和遇到的问题:

我先通过Selenium加载页面,点击「分段时间」标签,然后反复滚动页面确保所有内容都加载完成,总共抓取到了168个单元格。为了筛选出有效的分段时间,我设置了规则:只保留符合格式的小数(比如61.58或6.5),并且排除10及以下的数值——我以为这些是马匹编号(比如5)。最后处理下来只拿到了9匹马的数据,每匹补全8个分段时间(缺失的用0.00填充)。

预期结果:拿到88条分段时间,11匹马每匹对应8个有效数据,形成完整的数据集。
实际结果:只拿到77个有效单元格,仅处理了9匹马。

我收集到的调试信息:

All time cells found: 168
First 10 raw time cells: ['Pos', 'Silk', 'Horse', '1', '', ...]
Filtered time cells found: 77
Accepted samples: ['61.58', '56.55', '30.02', '27.18', '26.97']
Rejected samples: ['5.', '11.', '6.', '4.', '7.', '10.', '2.', '8.', '1.', '3.']
Horses Found in final data: 9

我的代码尝试:

from selenium import webdriver
from selenium.webdriver.firefox.service import Service
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
import pandas as pd
import time

# Firefox service
service = Service(executable_path="C:\\Scraping\\geckodriver.exe")
driver = webdriver.Firefox(service=service)

try:
    # Racecard URL
    url = "https://www.attheraces.com/racecard/Cheltenham/11-March-2025/1320"
    driver.get(url)
    time.sleep(15)  # Extended wait to allow full page load
    driver.save_screenshot("C:\\Scraping\\initial_load.png")
    print("Initial load captured")

    # Debug initial page source
    with open("C:\\Scraping\\page_source_initial.html", "w", encoding="utf-8") as f:
        f.write(driver.page_source)
    print("Initial page source saved")

    # Accept cookies
    try:
        cookie_button = WebDriverWait(driver, 10).until(
            EC.element_to_be_clickable((By.XPATH, '//button[text()="Accept All"]'))
        )
        cookie_button.click()
        print("Cookies accepted")
        time.sleep(5)
    except Exception as e:
        print("No cookie prompt found:", e)

    # Click Sectional Times tab with presence check and forced click
    print("Attempting to click Sectional Times tab...")
    try:
        tab = WebDriverWait(driver, 40).until(
            EC.presence_of_element_located((By.XPATH, '//a[contains(@class, "tab") and contains(normalize-space(.), "Sectional")]'))
        )
        driver.execute_script("arguments[0].click();", tab)
        print("Tab clicked")
    except Exception as e:
        print(f"Tab not found or click failed: {e}")
        driver.save_screenshot("C:\\Scraping\\tab_click_error.png")
    time.sleep(15)

    # Wait for tab container
    print("Waiting for tab container...")
    WebDriverWait(driver, 30).until(
        EC.presence_of_element_located((By.XPATH, "//div[@id='tab-sectionals-times']"))
    )
    print("Tab container found")

    # Multiple scroll attempts to load content
    print("Scrolling to load content...")
    for _ in range(5):
        driver.execute_script("window.scrollTo(0, document.body.scrollHeight);")
        time.sleep(5)
        driver.execute_script("window.scrollTo(0, 0);")
        time.sleep(2)
        driver.execute_script("window.scrollTo(0, document.body.scrollHeight);")
        time.sleep(5)
    driver.save_screenshot("C:\\Scraping\\after_scroll.png")
    print("Scroll completed")

    # Wait for sectional times with retry and extended timeout
    print("Waiting for sectional times...")
    max_attempts = 3
    for attempt in range(max_attempts):
        try:
            WebDriverWait(driver, 90).until(  # Increased to 90 seconds
                EC.presence_of_all_elements_located((By.XPATH, "//div[@id='tab-sectionals-times']//div[contains(@class, 'card-cell')]"))
            )
            print("Sectional times found")
            break
        except Exception as e:
            print(f"Attempt {attempt + 1} failed: {e}")
            if attempt < max_attempts - 1:
                print("Retrying...")
                time.sleep(10)
            else:
                print("Max retries reached, raising error")
                driver.save_screenshot("C:\\Scraping\\error_screenshot.png")
                raise
    driver.save_screenshot("C:\\Scraping\\after_times_wait.png")

    # Debug page source
    with open("C:\\Scraping\\page_source_after_scroll.html", "w", encoding="utf-8") as f:
        f.write(driver.page_source)

    # Headers
    headers = ["start-12f", "12f-8f", "8f-6f", "6f-4f", "4f-2f", "2f-1f", "1f-finish", "finish"]
    all_data = []

    # Race info
    race_name = driver.find_element(By.TAG_NAME, "h1").text.strip() if driver.find_elements(By.TAG_NAME, "h1") else "Unknown Race"
    distance = driver.find_element(By.CLASS_NAME, "p--large.font-weight--semibold").text.strip() if driver.find_elements(By.CLASS_NAME, "p--large.font-weight--semibold") else "0f"
    meeting = "cheltenham"

    # Scrape horse names
    print("Scraping horse names...")
    horse_elements = driver.find_elements(By.XPATH, "//div[contains(@class, 'card-entry')]//h2//a[contains(@class, 'horse__link')]")
    horses = [elem.text.strip() for elem in horse_elements if elem.text.strip() and elem.is_displayed()]
    print("Horses found:", len(horses))

    # Scrape sectional times with corrected alignment
    print("Scraping sectional times...")
    time_cells = driver.find_elements(By.XPATH, "//div[@id='tab-sectionals-times']//div[contains(@class, 'card-cell')]")
    print("All time cells found:", len(time_cells))
    print("First 10 raw time cells:", [cell.text.strip() for cell in time_cells[:10]])  # Debug raw data
    # Filter for time cells, accepting cleanable decimals after skipping headers
    time_cells_filtered = []
    accepted_samples = []
    rejected_samples = []
    header_count = 69  # Estimated based on raw data pattern
    for i, cell in enumerate(time_cells[header_count:]):  # Skip initial headers
        spans = cell.find_elements(By.XPATH, ".//span[contains(@class, 'visible')]")
        for span in spans:
            text = span.text.strip()
            cleaned = text.replace(' ', '').replace('x', '')
            if '.' in text and cleaned.replace('.', '').isdigit():
                parts = cleaned.split('.')
                if len(parts) == 2 and all(part.isdigit() for part in parts) and len(parts[0]) in [1, 2] and len(parts[1]) in [1, 2] and float(cleaned) > 10:
                    time_cells_filtered.append(cell)
                    accepted_samples.append(text)
                else:
                    rejected_samples.append(text)
    print("Filtered time cells found:", len(time_cells_filtered))
    if accepted_samples:
        print("Accepted samples:", accepted_samples[:5])  # Show first 5 accepted
    if rejected_samples:
        print("Rejected samples:", rejected_samples[:10])  # Show first 10 rejected

备注:内容来源于stack exchange,提问作者Audrey E Unique Expressions Ar

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.13 19:09:28