You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

新手求助:BeautifulSoup4爬取英超赛季射手数据遇JS过滤及解析问题

Hey there! Let's break down your problem and fix it step by step.

问题分析

You're facing two core issues here:

  • Getting historical total data instead of the target season's stats: The Premier League site loads season-specific stats via JavaScript dynamic rendering. When you use urlopen, you only grab the initial static HTML of the page—at this point, the 2017/18 season data hasn't been loaded by the browser's JS engine yet, so BeautifulSoup picks up the default historical total stats instead.
  • "Couldn't find a Tree Builder" error: This happens because the parser you specified (like lxml or html5lib) isn't installed on your system. BeautifulSoup needs these external parsers to process HTML properly.
解决方案

Let's tackle these issues one by one:

1. Fix the Parser Error

First, install a valid HTML parser. Run one of these commands in your terminal based on your preference:

# Install lxml (fast and popular)
pip install lxml

# Or install html5lib (more lenient with messy HTML)
pip install html5lib

After installation, your original script won't throw the parser error anymore—but we still need to solve the dynamic data problem.

2. Fetch JS-Rendered Target Season Data

Since the stats are loaded dynamically, you have two reliable ways to get the correct data:

Method 1: Use Selenium to Simulate a Browser (Beginner-Friendly)

Selenium mimics a real browser, letting you wait for JS to finish loading the data before scraping. Here's how to set it up:

  1. Install Selenium:
pip install selenium
  1. Download a browser driver (e.g., ChromeDriver for Google Chrome—make sure it matches your browser version) and place it in your system PATH or the same folder as your script.
  2. Run this modified script:
from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from bs4 import BeautifulSoup

# Initialize Chrome driver (omit executable_path if driver is in PATH)
driver = webdriver.Chrome(executable_path='./chromedriver')
stat_page = 'https://www.premierleague.com/stats/top/players/goals?se=79'
driver.get(stat_page)

# Wait up to 10 seconds for the stats table to load
wait = WebDriverWait(driver, 10)
wait.until(EC.presence_of_element_located((By.CLASS_NAME, 'statsTableContainer')))

# Grab the fully rendered page source
page_source = driver.page_source
driver.quit()

# Parse with BeautifulSoup
soup = BeautifulSoup(page_source, 'lxml')
stats = soup.find('tbody', attrs={'class': 'statsTableContainer'})
print(stats.text.strip())

Instead of scraping HTML, you can find the API that the site uses to load stats. Open your browser's DevTools (F12), go to the Network tab, refresh the page, and look for API requests. For your target season, the API endpoint looks like this:

import requests

# Add headers to mimic a real browser request
headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36",
    "Origin": "https://www.premierleague.com"
}

# API endpoint for 2017/18 season goals (se=79 maps to compSeasons=79)
api_url = "https://footballapi.pulselive.com/football/stats/top/players/goals?page=0&pageSize=20&compSeasons=79&compSeasonId=79&statId=21"
response = requests.get(api_url, headers=headers)
data = response.json()

# Extract and print player stats
for player in data['content']:
    player_name = player['owner']['name']['display']
    goals = player['value']
    print(f"{player_name}: {goals} goals")

This method is faster because you get structured JSON data directly, no HTML parsing needed.

Why Your Original Script Failed

When you use urlopen, you're only getting the raw static HTML sent by the server. The site's JS hasn't run yet to load the 2017/18 season data—so the <tbody> you're scraping contains the default historical stats instead of your target season's data.

内容的提问来源于stack exchange,提问作者daen23

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 06:35:54