You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Selenium与Beautiful Soup爬取雅虎每日幻想网页表格数据失败求助

Fixing Yahoo Daily Fantasy Table Scraping Issues

Let's break down what's going wrong with your code and get that table data extracted properly:

Key Problems in Your Current Code

  • Incorrect element locator: You're using By.TAG_NAME to target data-tst-player-id, but that's an attribute, not a tag name. This means your wait condition never actually finds the right element.
  • Forgetting parentheses on driver.quit: driver.quit without () doesn't execute the browser shutdown, which can leave processes hanging and cause unexpected behavior.
  • Passing a single element to BeautifulSoup: You're trying to parse a single web element instead of the full page source. BeautifulSoup needs the entire HTML of the page to locate all table elements.
  • Not waiting for the table specifically: Even if you found a player element, the table itself might still be loading in the background.

Corrected Code

Here's an updated version that addresses these issues:

from bs4 import BeautifulSoup
from selenium import webdriver
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait as wait

# Initialize the driver
driver = webdriver.Chrome()
driver.get("https://sports.yahoo.com/dailyfantasy/contest/5416455/setlineup")

try:
    # Wait for all player rows to load (adjust selector if needed)
    # Using CSS selector to target elements with data-tst-player-id attribute
    wait(driver, 15).until(
        EC.presence_of_all_elements_located((By.CSS_SELECTOR, "[data-tst-player-id]"))
    )
    
    # Get the full page source after elements are loaded
    page_source = driver.page_source
    soup = BeautifulSoup(page_source, 'lxml')
    
    # Find all rows with player data
    player_rows = soup.select("[data-tst-player-id]")
    
    # Write structured extracted data to your file
    with open('test.txt','w', encoding='utf-8') as f_out:
        for row in player_rows:
            # Extract specific data points (adjust based on what you need)
            player_name = row.find("div", class_="PlayerName__Name").text.strip() if row.find("div", class_="PlayerName__Name") else "N/A"
            position = row.find("div", class_="PlayerPosition").text.strip() if row.find("div", class_="PlayerPosition") else "N/A"
            f_out.write(f"Player: {player_name}, Position: {position}\n")
            
finally:
    # Always ensure the driver quits, even if an error occurs
    driver.quit()

Additional Notes

  • Dynamic page changes: Since the URL might update weekly, you'll need to periodically check if the CSS selectors (like [data-tst-player-id], PlayerName__Name) are still valid. Yahoo might adjust their class names or attributes over time.
  • Wait time adjustments: If the page loads slowly, increase the wait time from 15 seconds to something more appropriate.
  • Authentication: If this page requires you to log in to view the lineup, add code to handle login (using driver.find_element to enter credentials and click the login button).

内容的提问来源于stack exchange,提问作者user3362580

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.14 08:06:37