You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Pandas无法读取网页中存在的Advanced Receiving表格问题求助

问题:读取Pro-Football-Reference的Advanced Receiving表格报错

我在Jupyter Lab中编写代码,尝试让pandas像读取“Advanced Passing”表格一样读取“Advanced Receiving”表格,但持续报错:ValueError: No tables found matching pattern 'Advanced Receiving'。目标网页链接:https://www.pro-football-reference.com/teams/buf/2023_advanced.htm。我已检查动态内容加载问题和表格逻辑,但问题仍未解决。

我的代码如下:

import requests
from bs4 import BeautifulSoup


standings_url = "https://www.pro-football-reference.com/years/2023/index.htm"



data = requests.get(standings_url)
soup = BeautifulSoup(data.text)

standings_table1 = soup.select('table.stats_table')[0]
standings_table2 = soup.select('table.stats_table')[1]
links1 = standings_table1.find_all('a')
links2 = standings_table2.find_all('a')
all_links = [l.get('href') for l in links1] + [l.get('href') for l in links2]
team_links = [l for l in all_links if '/teams/' in l]

base_url = "https://www.pro-football-reference.com"
full_team_links = [base_url + l for l in all_links]
full_team_links

team_url = full_team_links[0]
data = requests.get(team_url)
data.text

import pandas as pd
games = pd.read_html(data.text, match="Schedule & Game Results")
games[0]

soup = BeautifulSoup(data.text, 'html5lib')
links = soup.find_all('a')
links =[l.get("href") for l in links]
links = [l for l in links if l and '/2023_advanced.htm' in l]
links

data = requests.get(f"https://www.pro-football-reference.com{links[0]}")
data.text

passing = pd.read_html(data.text, match="Advanced Passing")[0]
receiving = pd.read_html(data.text, match="Advanced Receiving")[0]
passing.head()
receiving.head()

问题原因

Pro-Football-Reference网站的“Advanced Receiving”表格被包裹在HTML注释(<!--和-->)中,而pd.read_html默认只会解析页面中直接可见的HTML元素,无法识别注释内的表格内容。

解决方法

方法1:移除HTML注释标签

获取页面内容后,直接替换掉注释标签,让表格内容成为可解析的HTML:

# 替换原代码中读取advanced页面后的部分
data = requests.get(f"https://www.pro-football-reference.com{links[0]}")
# 移除页面中的HTML注释标签
processed_html = data.text.replace('<!--', '').replace('-->', '')

# 现在可以正常读取两个表格
passing = pd.read_html(processed_html, match="Advanced Passing")[0]
receiving = pd.read_html(processed_html, match="Advanced Receiving")[0]
passing.head()
receiving.head()

方法2:用BeautifulSoup提取注释中的表格

精准定位包含目标表格的注释,再单独解析:

from bs4 import Comment

# 替换原代码中读取advanced页面后的部分
data = requests.get(f"https://www.pro-football-reference.com{links[0]}")
soup = BeautifulSoup(data.text, 'html5lib')

# 遍历所有HTML注释,找到包含Advanced Receiving的内容
receiving_comment = None
for comment in soup.find_all(string=lambda text: isinstance(text, Comment)):
    if 'Advanced Receiving' in comment:
        receiving_comment = comment
        break

# 读取表格
passing = pd.read_html(data.text, match="Advanced Passing")[0]
receiving = pd.read_html(receiving_comment)[0]
passing.head()
receiving.head()

内容的提问来源于stack exchange,提问作者Sloh

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.17 12:45:02