使用BeautifulSoup提取match-info类a标签href无结果,求解决方案
解决Understat网页match-info链接提取失败的思路
1. 补全缺失的依赖库导入
你的代码里调用了requests.get()但未导入requests库,这会直接导致请求失败,先补上导入语句:
import requests # 新增这行 import warnings import numpy as np from datetime import datetime import json from bs4 import BeautifulSoup warnings.filterwarnings('ignore') url = "https://understat.com/league/EPL/2022" response = requests.get(url) # ... 后续代码不变
2. 处理动态渲染的页面内容
Understat的比赛链接是通过JavaScript动态渲染到页面的,直接用BeautifulSoup抓取静态HTML无法获取到这些元素,有两种可行方案:
方案A:直接解析页面内嵌的JSON数据
页面源代码中包含window.__INITIAL_STATE__变量,里面存储了所有比赛的完整数据,直接提取并解析这个JSON就能拿到match的href:
import requests import json from bs4 import BeautifulSoup url = "https://understat.com/league/EPL/2022" response = requests.get(url) soup = BeautifulSoup(response.content, "html.parser") # 提取页面中的INITIAL_STATE数据 script_tag = soup.find('script', text=lambda t: '__INITIAL_STATE__' in t) raw_data = script_tag.text.split('window.__INITIAL_STATE__ = ')[1].rstrip(';') data = json.loads(raw_data) # 遍历所有比赛,提取href for match in data['dates']: if 'matches' in match: for m in match['matches']: print(f"match/{m['id']}")
方案B:使用浏览器渲染工具加载页面
用Selenium模拟浏览器加载页面,等待JavaScript渲染完成后再抓取DOM:
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC url = "https://understat.com/league/EPL/2022" driver = webdriver.Chrome() # 需要提前安装对应版本的ChromeDriver driver.get(url) # 等待元素加载完成 WebDriverWait(driver, 10).until( EC.presence_of_all_elements_located((By.CLASS_NAME, "match-info")) ) # 提取所有match-info的href links = driver.find_elements(By.CLASS_NAME, "match-info") for link in links: href = link.get_attribute("href") print(href.split('understat.com/')[-1]) # 只输出match/xxx部分 driver.quit()
3. 优化选择器精度
可以通过同时指定class和data-isresult属性来缩小选择范围,确保只匹配目标元素:
# 修改find_all的参数,同时过滤class和data-isresult属性 for link in soup.find_all("a", {"class": "match-info", "data-isresult": "true"}): href = link.get("href") print(href)
注:这个方法仅在静态HTML中存在该元素时有效,若元素是动态渲染的,还是需要用前面的动态渲染处理方案。
4. 添加请求头模拟浏览器访问
部分网站会拦截无标识的请求,添加User-Agent等请求头可以避免被拦截:
headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36' } response = requests.get(url, headers=headers)
内容的提问来源于stack exchange,提问作者Paul Corcoran
相关产品推荐
相关产品推荐

