You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用Python的BeautifulSoup爬取“show more”按钮后的英超赛事数据?

How to Scrape Hidden Premier League Results (After "Show More" Button) with BeautifulSoup?

I'm using Python's BeautifulSoup to scrape football match data from the 2020-21 Premier League results page on Sky Sports. The site only shows the first 200 matches, and the remaining 180 require clicking a "Show More" button to view. Since clicking the button doesn't change the URL, I can't just modify the URL to get the rest of the content. My current code only fetches the first 200 matches—how do I get the HTML content after the "Show More" button is clicked?

My current code:

from bs4 import BeautifulSoup
import requests
scores_html_text = requests.get('https://www.skysports.com/premier-league-results/2020-21').text
scores_soup = BeautifulSoup(scores_html_text, 'lxml')
fixtures = scores_soup.find_all('div', class_ = 'fixres__item')

Hey there! The problem here is that the extra matches are loaded dynamically with JavaScript—the requests library only pulls the initial static HTML, so it can't access content that loads after user interactions like clicking a button. Below are two reliable solutions to grab all 380 matches:

1. Use Selenium to Simulate Browser Actions

This is the most straightforward approach, especially if you're new to dynamic scraping. Selenium controls a real browser, so it can click the "Show More" button and wait for new content to load.

Steps & Code:

First, install Selenium and download the appropriate browser driver (e.g., ChromeDriver for Google Chrome, matching your browser version):

pip install selenium

Then use this code:

from bs4 import BeautifulSoup
from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
import time

# Initialize Chrome browser (replace with your preferred browser driver)
driver = webdriver.Chrome()
url = "https://www.skysports.com/premier-league-results/2020-21"
driver.get(url)

# Keep clicking "Show More" until the button disappears
while True:
    try:
        # Wait up to 10 seconds for the button to become clickable
        show_more_btn = WebDriverWait(driver, 10).until(
            EC.element_to_be_clickable((By.CLASS_NAME, "ss-btn__text"))
        )
        show_more_btn.click()
        # Give the page time to load new content (adjust if needed)
        time.sleep(2)
    except:
        # Exit loop when the button is no longer found
        break

# Grab the full, dynamically loaded page source
full_html = driver.page_source
driver.quit()

# Parse with BeautifulSoup as usual
scores_soup = BeautifulSoup(full_html, 'lxml')
fixtures = scores_soup.find_all('div', class_='fixres__item')

# Verify you have all matches
print(f"Successfully scraped {len(fixtures)} matches!")

Pros & Cons:

  • ✅ Pros: No need to analyze backend API requests; works exactly like a human user.
  • ❌ Cons: Slower than API scraping; requires a browser driver; uses more system resources.

2. Simulate the AJAX Request (Faster, No Browser Needed)

When you click "Show More", the site sends an AJAX request to its backend to fetch additional matches. You can replicate this request directly with requests to get the new content without a browser.

Steps & Code:

  1. Open your browser's DevTools (F12), go to the Network tab, and click "Show More" to capture the AJAX request.
  2. Copy the request URL, headers, and form data (you'll need these for your code).

Here's an example of how to implement this:

from bs4 import BeautifulSoup
import requests
import time

base_url = "https://www.skysports.com/premier-league-results/2020-21"
# Replace this with the actual AJAX endpoint you found in DevTools
ajax_endpoint = "https://www.skysports.com/a/ajax/service"

# Create a session to persist cookies/headers between requests
session = requests.Session()
initial_response = session.get(base_url)
initial_soup = BeautifulSoup(initial_response.text, 'lxml')

# Grab initial 200 matches
fixtures = initial_soup.find_all('div', class_='fixres__item')

# Set up headers and data for AJAX requests (adjust based on DevTools capture)
headers = {
    "X-Requested-With": "XMLHttpRequest",
    "Content-Type": "application/x-www-form-urlencoded"
}
offset = 200  # Start after the first 200 matches
limit = 100   # Number of matches loaded per "Show More" click

while True:
    # Form data from DevTools (adjust parameters as needed)
    data = {
        "service": "fixres",
        "action": "getMoreFixtures",
        "offset": str(offset),
        "limit": str(limit),
        "competition": "premier-league",
        "season": "2020-21",
        # Grab CSRF token from the initial page (required for most sites)
        "csrfToken": initial_soup.find("meta", attrs={"name": "csrf-token"})["content"]
    }

    # Send the AJAX request
    ajax_response = session.post(ajax_endpoint, headers=headers, data=data)
    
    # Exit loop if no more content is returned
    if not ajax_response.text:
        break
    
    # Parse the returned HTML fragment
    ajax_soup = BeautifulSoup(ajax_response.text, 'lxml')
    new_fixtures = ajax_soup.find_all('div', class_='fixres__item')
    
    if not new_fixtures:
        break
    
    # Add new matches to our list
    fixtures.extend(new_fixtures)
    offset += limit
    # Avoid overwhelming the server with requests
    time.sleep(1)

print(f"Successfully scraped {len(fixtures)} matches!")

Pros & Cons:

  • ✅ Pros: Much faster; uses minimal system resources; no browser required.
  • ❌ Cons: Requires inspecting the site's backend requests; may break if the site updates its API parameters.

内容的提问来源于stack exchange,提问作者Jin

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.30 10:17:53