Pandas无法读取网页中存在的Advanced Receiving表格问题求助
问题:读取Pro-Football-Reference的Advanced Receiving表格报错
我在Jupyter Lab中编写代码,尝试让pandas像读取“Advanced Passing”表格一样读取“Advanced Receiving”表格,但持续报错:ValueError: No tables found matching pattern 'Advanced Receiving'。目标网页链接:https://www.pro-football-reference.com/teams/buf/2023_advanced.htm。我已检查动态内容加载问题和表格逻辑,但问题仍未解决。
我的代码如下:
import requests from bs4 import BeautifulSoup standings_url = "https://www.pro-football-reference.com/years/2023/index.htm" data = requests.get(standings_url) soup = BeautifulSoup(data.text) standings_table1 = soup.select('table.stats_table')[0] standings_table2 = soup.select('table.stats_table')[1] links1 = standings_table1.find_all('a') links2 = standings_table2.find_all('a') all_links = [l.get('href') for l in links1] + [l.get('href') for l in links2] team_links = [l for l in all_links if '/teams/' in l] base_url = "https://www.pro-football-reference.com" full_team_links = [base_url + l for l in all_links] full_team_links team_url = full_team_links[0] data = requests.get(team_url) data.text import pandas as pd games = pd.read_html(data.text, match="Schedule & Game Results") games[0] soup = BeautifulSoup(data.text, 'html5lib') links = soup.find_all('a') links =[l.get("href") for l in links] links = [l for l in links if l and '/2023_advanced.htm' in l] links data = requests.get(f"https://www.pro-football-reference.com{links[0]}") data.text passing = pd.read_html(data.text, match="Advanced Passing")[0] receiving = pd.read_html(data.text, match="Advanced Receiving")[0] passing.head() receiving.head()
问题原因
Pro-Football-Reference网站的“Advanced Receiving”表格被包裹在HTML注释(<!--和-->)中,而pd.read_html默认只会解析页面中直接可见的HTML元素,无法识别注释内的表格内容。
解决方法
方法1:移除HTML注释标签
获取页面内容后,直接替换掉注释标签,让表格内容成为可解析的HTML:
# 替换原代码中读取advanced页面后的部分 data = requests.get(f"https://www.pro-football-reference.com{links[0]}") # 移除页面中的HTML注释标签 processed_html = data.text.replace('<!--', '').replace('-->', '') # 现在可以正常读取两个表格 passing = pd.read_html(processed_html, match="Advanced Passing")[0] receiving = pd.read_html(processed_html, match="Advanced Receiving")[0] passing.head() receiving.head()
方法2:用BeautifulSoup提取注释中的表格
精准定位包含目标表格的注释,再单独解析:
from bs4 import Comment # 替换原代码中读取advanced页面后的部分 data = requests.get(f"https://www.pro-football-reference.com{links[0]}") soup = BeautifulSoup(data.text, 'html5lib') # 遍历所有HTML注释,找到包含Advanced Receiving的内容 receiving_comment = None for comment in soup.find_all(string=lambda text: isinstance(text, Comment)): if 'Advanced Receiving' in comment: receiving_comment = comment break # 读取表格 passing = pd.read_html(data.text, match="Advanced Passing")[0] receiving = pd.read_html(receiving_comment)[0] passing.head() receiving.head()
内容的提问来源于stack exchange,提问作者Sloh
相关产品推荐
相关产品推荐

