如何使用Python的BeautifulSoup爬取“show more”按钮后的英超赛事数据?
I'm using Python's BeautifulSoup to scrape football match data from the 2020-21 Premier League results page on Sky Sports. The site only shows the first 200 matches, and the remaining 180 require clicking a "Show More" button to view. Since clicking the button doesn't change the URL, I can't just modify the URL to get the rest of the content. My current code only fetches the first 200 matches—how do I get the HTML content after the "Show More" button is clicked?
My current code:
from bs4 import BeautifulSoup import requests scores_html_text = requests.get('https://www.skysports.com/premier-league-results/2020-21').text scores_soup = BeautifulSoup(scores_html_text, 'lxml') fixtures = scores_soup.find_all('div', class_ = 'fixres__item')
Hey there! The problem here is that the extra matches are loaded dynamically with JavaScript—the requests library only pulls the initial static HTML, so it can't access content that loads after user interactions like clicking a button. Below are two reliable solutions to grab all 380 matches:
1. Use Selenium to Simulate Browser Actions
This is the most straightforward approach, especially if you're new to dynamic scraping. Selenium controls a real browser, so it can click the "Show More" button and wait for new content to load.
Steps & Code:
First, install Selenium and download the appropriate browser driver (e.g., ChromeDriver for Google Chrome, matching your browser version):
pip install selenium
Then use this code:
from bs4 import BeautifulSoup from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC import time # Initialize Chrome browser (replace with your preferred browser driver) driver = webdriver.Chrome() url = "https://www.skysports.com/premier-league-results/2020-21" driver.get(url) # Keep clicking "Show More" until the button disappears while True: try: # Wait up to 10 seconds for the button to become clickable show_more_btn = WebDriverWait(driver, 10).until( EC.element_to_be_clickable((By.CLASS_NAME, "ss-btn__text")) ) show_more_btn.click() # Give the page time to load new content (adjust if needed) time.sleep(2) except: # Exit loop when the button is no longer found break # Grab the full, dynamically loaded page source full_html = driver.page_source driver.quit() # Parse with BeautifulSoup as usual scores_soup = BeautifulSoup(full_html, 'lxml') fixtures = scores_soup.find_all('div', class_='fixres__item') # Verify you have all matches print(f"Successfully scraped {len(fixtures)} matches!")
Pros & Cons:
- ✅ Pros: No need to analyze backend API requests; works exactly like a human user.
- ❌ Cons: Slower than API scraping; requires a browser driver; uses more system resources.
2. Simulate the AJAX Request (Faster, No Browser Needed)
When you click "Show More", the site sends an AJAX request to its backend to fetch additional matches. You can replicate this request directly with requests to get the new content without a browser.
Steps & Code:
- Open your browser's DevTools (F12), go to the Network tab, and click "Show More" to capture the AJAX request.
- Copy the request URL, headers, and form data (you'll need these for your code).
Here's an example of how to implement this:
from bs4 import BeautifulSoup import requests import time base_url = "https://www.skysports.com/premier-league-results/2020-21" # Replace this with the actual AJAX endpoint you found in DevTools ajax_endpoint = "https://www.skysports.com/a/ajax/service" # Create a session to persist cookies/headers between requests session = requests.Session() initial_response = session.get(base_url) initial_soup = BeautifulSoup(initial_response.text, 'lxml') # Grab initial 200 matches fixtures = initial_soup.find_all('div', class_='fixres__item') # Set up headers and data for AJAX requests (adjust based on DevTools capture) headers = { "X-Requested-With": "XMLHttpRequest", "Content-Type": "application/x-www-form-urlencoded" } offset = 200 # Start after the first 200 matches limit = 100 # Number of matches loaded per "Show More" click while True: # Form data from DevTools (adjust parameters as needed) data = { "service": "fixres", "action": "getMoreFixtures", "offset": str(offset), "limit": str(limit), "competition": "premier-league", "season": "2020-21", # Grab CSRF token from the initial page (required for most sites) "csrfToken": initial_soup.find("meta", attrs={"name": "csrf-token"})["content"] } # Send the AJAX request ajax_response = session.post(ajax_endpoint, headers=headers, data=data) # Exit loop if no more content is returned if not ajax_response.text: break # Parse the returned HTML fragment ajax_soup = BeautifulSoup(ajax_response.text, 'lxml') new_fixtures = ajax_soup.find_all('div', class_='fixres__item') if not new_fixtures: break # Add new matches to our list fixtures.extend(new_fixtures) offset += limit # Avoid overwhelming the server with requests time.sleep(1) print(f"Successfully scraped {len(fixtures)} matches!")
Pros & Cons:
- ✅ Pros: Much faster; uses minimal system resources; no browser required.
- ❌ Cons: Requires inspecting the site's backend requests; may break if the site updates its API parameters.
内容的提问来源于stack exchange,提问作者Jin

