Python爬取IPL海得拉巴日出队阵容:下拉年份筛选数据难题
Got it, let's tackle this problem. The core issue here is that when you select a year from the dropdown, the player list updates dynamically via JavaScript (AJAX)—the page doesn't reload or change URLs, so your initial BeautifulSoup scrape only gets the first loaded season's data. Here are two reliable approaches to fetch player data for every season:
Approach 1: Use Selenium to Simulate User Interaction
This method mimics how a real user would interact with the page—clicking each year option, waiting for the page to update, then scraping the new content. It's straightforward and works well if you're not comfortable digging into API requests.
Updated Code
from selenium import webdriver from selenium.webdriver.chrome.service import Service from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from bs4 import BeautifulSoup import time squad_url = "https://www.iplt20.com/teams/sunrisers-hyderabad/squad" # Initialize Chrome driver (Selenium 4+ compatible) service = Service("./chromedriver.exe") driver = webdriver.Chrome(service=service) driver.get(squad_url) all_season_data = {} try: # Wait for the dropdown trigger to load and click it to open the year list dropdown_trigger = WebDriverWait(driver, 10).until( EC.element_to_be_clickable((By.CLASS_NAME, "js-dropdown-trigger")) ) dropdown_trigger.click() # Get all year option elements year_options = WebDriverWait(driver, 10).until( EC.presence_of_all_elements_located((By.CLASS_NAME, "drop-down__dropdown-list__option")) ) for option in year_options: year_text = option.text # Click the year option option.click() # Wait for the page to update (verify the selected year matches) WebDriverWait(driver, 10).until( EC.text_to_be_present_in_element((By.CLASS_NAME, "js-drop-down-current"), year_text) ) time.sleep(1) # Add a short buffer to ensure player cards load # Parse the updated page source soup = BeautifulSoup(driver.page_source, "html.parser") player_cards = soup.find_all("div", class_="large-squad-list__player-card") # Extract player details season_players = [] for card in player_cards: player = { "name": card.find("h3", class_="large-squad-list__player-name").text.strip() if card.find("h3", class_="large-squad-list__player-name") else "N/A", "position": card.find("span", class_="large-squad-list__player-position").text.strip() if card.find("span", class_="large-squad-list__player-position") else "N/A", "nationality": card.find("span", class_="large-squad-list__player-country").text.strip() if card.find("span", class_="large-squad-list__player-country") else "N/A" } season_players.append(player) all_season_data[year_text] = season_players print(f"✅ Fetched {len(season_players)} players for {year_text}") # Re-open the dropdown for the next iteration dropdown_trigger.click() time.sleep(0.5) finally: driver.quit() # Example: Print data for each season for year, players in all_season_data.items(): print(f"\n--- {year} Season ---") for p in players: print(f"{p['name']} | {p['position']} | {p['nationality']}")
Key Notes:
- We use
WebDriverWaitto ensure elements are loaded before interacting with them (avoids race conditions). - After clicking a year, we wait until the selected year text updates in the dropdown to confirm the page has refreshed.
- We re-open the dropdown after each selection since it closes automatically when you click an option.
Approach 2: Directly Call the AJAX API (Faster & More Efficient)
Instead of simulating clicks, you can bypass the browser entirely by calling the same API the page uses to load player data. This is faster and uses fewer resources.
How to Find the API:
- Open your browser's DevTools (F12) → Go to the Network tab.
- Select a year from the dropdown—you'll see a new request appear (look for something like
squadin the request name). - Check the request URL and parameters—you'll notice it uses the
data-optionvalues (e.g.,ipl2020) as a season parameter.
API-Based Code
import requests # Base API URL (found via browser DevTools) base_api = "https://www.iplt20.com/api/teams/squad" team_slug = "sunrisers-hyderabad" # List of season codes (matches the data-option values from the dropdown) seasons = ["ipl2020", "ipl2019", "ipl2018", "ipl2017", "ipl2016", "ipl2015", "ipl2014", "ipl2013", "ipl2012", "ipl2011", "ipl2010", "ipl2009", "ipl2008"] all_season_data = {} for season_code in seasons: params = { "team": team_slug, "season": season_code } response = requests.get(base_api, params=params) if response.status_code == 200: season_year = season_code.replace("ipl", "") data = response.json() # Extract player data from the JSON response season_players = [] for player in data["players"]: season_players.append({ "name": player["fullName"], "position": player["position"], "nationality": player["country"], "dob": player["dateOfBirth"] }) all_season_data[season_year] = season_players print(f"✅ Fetched {len(season_players)} players for {season_year}") else: print(f"❌ Failed to fetch data for {season_code.replace('ipl', '')} (Status code: {response.status_code})") # Example: Print data for year, players in all_season_data.items(): print(f"\n--- {year} Season ---") for p in players: print(f"{p['name']} | {p['position']} | {p['nationality']}")
Key Notes:
- This method is much faster since it doesn't require a browser to load the entire page.
- The API returns structured JSON data, so you don't have to parse HTML—just extract the fields you need directly.
- Keep in mind that APIs can change over time, so you may need to re-check the request parameters if this stops working.
内容的提问来源于stack exchange,提问作者lokesh kumar

