Python爬取FanDuel页面href链接遇阻(class动态变化)
Hey there! Let's work through this problem step by step—you're on the right track using the fixed data-test-id attribute, but there are a couple of key issues getting in your way.
1. Fix the Basic Code Error
Your line game = soup.find_all('a', {'data-test-id':"ContestCardEnterLink"})[('href')] is throwing an error because find_all() returns a list of matching elements, not a single tag. You can't directly index a list with ['href']. Here's how to fix that first:
import requests from bs4 import BeautifulSoup url = "https://www.fanduel.com/contests/mlb/96" # Keep your auth cookie if it's required for access site = requests.get(url, cookies={'X-Auth-Token':'MY TOKEN IS HERE'}) soup = BeautifulSoup(site.content, 'lxml') # Get all matching <a> tags with the target data-test-id game_links = soup.find_all('a', {'data-test-id':"ContestCardEnterLink"}) # Check if we found links, then safely extract hrefs if game_links: for link in game_links: # Use .get('href') to avoid AttributeError if the attribute is missing href = link.get('href') print(href) else: print("No matching links found in the static HTML.")
2. Address Dynamic Content (The Likely Main Problem)
If the fixed code still returns nothing, that's almost certainly because FanDuel loads most contest content dynamically with JavaScript. The requests library only fetches the initial static HTML, not the content that loads after the page runs its JS scripts.
To handle this, you'll need a tool that can render the full page like a browser. Selenium is a popular choice for this scenario. Here's a working example:
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC url = "https://www.fanduel.com/contests/mlb/96" # Initialize Chrome (make sure you have ChromeDriver installed and in your system PATH) driver = webdriver.Chrome() driver.get(url) try: # Wait up to 10 seconds for the target links to load (adjust the timeout if needed) game_links = WebDriverWait(driver, 10).until( EC.presence_of_all_elements_located((By.CSS_SELECTOR, 'a[data-test-id="ContestCardEnterLink"]')) ) # Extract and print each href attribute for link in game_links: href = link.get_attribute('href') print(href) finally: # Always close the browser when finished to avoid hanging processes driver.quit()
3. Quick Notes on Authentication
If the page requires you to be logged in to see contests, you can either:
- Log in manually through the Selenium browser window before the extraction code runs, or
- Export your logged-in cookies from your regular browser and import them into Selenium to skip manual login.
内容的提问来源于stack exchange,提问作者sbiondio

