新手求助:BeautifulSoup4爬取英超赛季射手数据遇JS过滤及解析问题
Hey there! Let's break down your problem and fix it step by step.
You're facing two core issues here:
- Getting historical total data instead of the target season's stats: The Premier League site loads season-specific stats via JavaScript dynamic rendering. When you use
urlopen, you only grab the initial static HTML of the page—at this point, the 2017/18 season data hasn't been loaded by the browser's JS engine yet, so BeautifulSoup picks up the default historical total stats instead. - "Couldn't find a Tree Builder" error: This happens because the parser you specified (like
lxmlorhtml5lib) isn't installed on your system. BeautifulSoup needs these external parsers to process HTML properly.
Let's tackle these issues one by one:
1. Fix the Parser Error
First, install a valid HTML parser. Run one of these commands in your terminal based on your preference:
# Install lxml (fast and popular) pip install lxml # Or install html5lib (more lenient with messy HTML) pip install html5lib
After installation, your original script won't throw the parser error anymore—but we still need to solve the dynamic data problem.
2. Fetch JS-Rendered Target Season Data
Since the stats are loaded dynamically, you have two reliable ways to get the correct data:
Method 1: Use Selenium to Simulate a Browser (Beginner-Friendly)
Selenium mimics a real browser, letting you wait for JS to finish loading the data before scraping. Here's how to set it up:
- Install Selenium:
pip install selenium
- Download a browser driver (e.g., ChromeDriver for Google Chrome—make sure it matches your browser version) and place it in your system PATH or the same folder as your script.
- Run this modified script:
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from bs4 import BeautifulSoup # Initialize Chrome driver (omit executable_path if driver is in PATH) driver = webdriver.Chrome(executable_path='./chromedriver') stat_page = 'https://www.premierleague.com/stats/top/players/goals?se=79' driver.get(stat_page) # Wait up to 10 seconds for the stats table to load wait = WebDriverWait(driver, 10) wait.until(EC.presence_of_element_located((By.CLASS_NAME, 'statsTableContainer'))) # Grab the fully rendered page source page_source = driver.page_source driver.quit() # Parse with BeautifulSoup soup = BeautifulSoup(page_source, 'lxml') stats = soup.find('tbody', attrs={'class': 'statsTableContainer'}) print(stats.text.strip())
Method 2: Call the Direct API (More Efficient, Recommended)
Instead of scraping HTML, you can find the API that the site uses to load stats. Open your browser's DevTools (F12), go to the Network tab, refresh the page, and look for API requests. For your target season, the API endpoint looks like this:
import requests # Add headers to mimic a real browser request headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36", "Origin": "https://www.premierleague.com" } # API endpoint for 2017/18 season goals (se=79 maps to compSeasons=79) api_url = "https://footballapi.pulselive.com/football/stats/top/players/goals?page=0&pageSize=20&compSeasons=79&compSeasonId=79&statId=21" response = requests.get(api_url, headers=headers) data = response.json() # Extract and print player stats for player in data['content']: player_name = player['owner']['name']['display'] goals = player['value'] print(f"{player_name}: {goals} goals")
This method is faster because you get structured JSON data directly, no HTML parsing needed.
When you use urlopen, you're only getting the raw static HTML sent by the server. The site's JS hasn't run yet to load the 2017/18 season data—so the <tbody> you're scraping contains the default historical stats instead of your target season's data.
内容的提问来源于stack exchange,提问作者daen23

