使用Selenium与Beautiful Soup爬取雅虎每日幻想网页表格数据失败求助
Fixing Yahoo Daily Fantasy Table Scraping Issues
Let's break down what's going wrong with your code and get that table data extracted properly:
Key Problems in Your Current Code
- Incorrect element locator: You're using
By.TAG_NAMEto targetdata-tst-player-id, but that's an attribute, not a tag name. This means your wait condition never actually finds the right element. - Forgetting parentheses on
driver.quit:driver.quitwithout()doesn't execute the browser shutdown, which can leave processes hanging and cause unexpected behavior. - Passing a single element to BeautifulSoup: You're trying to parse a single web element instead of the full page source. BeautifulSoup needs the entire HTML of the page to locate all table elements.
- Not waiting for the table specifically: Even if you found a player element, the table itself might still be loading in the background.
Corrected Code
Here's an updated version that addresses these issues:
from bs4 import BeautifulSoup from selenium import webdriver from selenium.webdriver.support import expected_conditions as EC from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait as wait # Initialize the driver driver = webdriver.Chrome() driver.get("https://sports.yahoo.com/dailyfantasy/contest/5416455/setlineup") try: # Wait for all player rows to load (adjust selector if needed) # Using CSS selector to target elements with data-tst-player-id attribute wait(driver, 15).until( EC.presence_of_all_elements_located((By.CSS_SELECTOR, "[data-tst-player-id]")) ) # Get the full page source after elements are loaded page_source = driver.page_source soup = BeautifulSoup(page_source, 'lxml') # Find all rows with player data player_rows = soup.select("[data-tst-player-id]") # Write structured extracted data to your file with open('test.txt','w', encoding='utf-8') as f_out: for row in player_rows: # Extract specific data points (adjust based on what you need) player_name = row.find("div", class_="PlayerName__Name").text.strip() if row.find("div", class_="PlayerName__Name") else "N/A" position = row.find("div", class_="PlayerPosition").text.strip() if row.find("div", class_="PlayerPosition") else "N/A" f_out.write(f"Player: {player_name}, Position: {position}\n") finally: # Always ensure the driver quits, even if an error occurs driver.quit()
Additional Notes
- Dynamic page changes: Since the URL might update weekly, you'll need to periodically check if the CSS selectors (like
[data-tst-player-id],PlayerName__Name) are still valid. Yahoo might adjust their class names or attributes over time. - Wait time adjustments: If the page loads slowly, increase the wait time from 15 seconds to something more appropriate.
- Authentication: If this page requires you to log in to view the lineup, add code to handle login (using
driver.find_elementto enter credentials and click the login button).
内容的提问来源于stack exchange,提问作者user3362580
相关产品推荐
相关产品推荐

