使用BeautifulSoup爬取Basketball Reference时遇NoneType属性错误求助
问题描述
我正在为Basketball Reference网站开发数据爬虫,用于提取指定NBA球员“Per Game”表格的表头与统计数据,并展示在自建页面中。相关代码如下:
@app.route('/', methods=['GET', 'POST']) def index(): if request.method == 'POST': player_name = request.form['player'] player_url = get_player_url(player_name) response = requests.get(player_url) soup = BeautifulSoup(response.content, 'html.parser') stats = soup.find('div', {'id':'all_per_game'}) table = stats.find('table') table_headings = [th.get_text() for th in table.find('tr').find_all('th')] table_data = [] for row in table.find_all('tr')[1:]: row_data = [td.get_text() for td in row.find_all('td')] table_data.append(row_data) return render_template('stats.html', player_name=player_name, headings=table_headings, data=table_data) else: return render_template('index.html', players=players)
运行本地Python应用时,在获取表格表头步骤出现如下错误:
File ¨main.py¨, line 48, in index table_headings = [th.get_text() for th in table.find('tr').find_all('th')] AttributeError: 'NoneType' object has no attribute find
我是BeautifulSoup新手,已查阅网络资料但未找到适配方案,请问如何正确提取表头?该问题是否与table对象类型有关?
解决方案
这个错误的核心是table或table.find('tr')返回了None——也就是BeautifulSoup没找到对应的元素,大概率是Basketball Reference的页面结构有特殊处理,或是你的选择器/请求方式有问题。
问题排查与修复步骤
先确认元素是否存在,添加调试与异常处理
先在代码中加入判断,避免None对象调用方法:stats = soup.find('div', {'id':'all_per_game'}) # 先判断stats是否存在,再找表格 table = stats.find('table') if stats else None # 如果table不存在,直接返回错误提示 if not table: return render_template('error.html', message='无法找到目标数据表格')如果
stats是None,要么是请求被网站拦截,要么是页面结构变更;如果stats存在但table是None,那大概率是表格被藏在了HTML注释里——Basketball Reference常用这种方式防爬虫。处理藏在注释中的表格
如果表格在注释内,需要先提取注释内容再解析:from bs4 import Comment stats = soup.find('div', {'id':'all_per_game'}) table = None if stats: # 先尝试直接找表格 table = stats.find('table') # 找不到的话遍历所有注释内容 if not table: comments = stats.find_all(string=lambda text: isinstance(text, Comment)) for comment in comments: comment_soup = BeautifulSoup(comment, 'html.parser') table = comment_soup.find('table') if table: break正确提取表头
不要直接用table.find('tr'),优先找thead标签(标准表格结构中表头会放在这里),避免误取数据行:table_headings = [] # 优先从thead提取表头 thead = table.find('thead') if thead: header_row = thead.find('tr') table_headings = [th.get_text(strip=True) for th in header_row.find_all('th')] else: # 没有thead的话,取表格第一个tr作为表头 header_row = table.find('tr') if header_row: table_headings = [th.get_text(strip=True) for th in header_row.find_all('th')]用
strip=True可以自动去除文本前后的空格和换行,让表头更整洁。添加请求头模拟浏览器
很多网站会拦截无标识的爬虫请求,给requests.get加上浏览器请求头:headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36' } response = requests.get(player_url, headers=headers)
完整修正后的核心代码片段
@app.route('/', methods=['GET', 'POST']) def index(): if request.method == 'POST': player_name = request.form['player'] player_url = get_player_url(player_name) # 添加浏览器请求头 headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36' } response = requests.get(player_url, headers=headers) soup = BeautifulSoup(response.content, 'html.parser') stats = soup.find('div', {'id':'all_per_game'}) table = None if stats: table = stats.find('table') # 从注释中找表格 if not table: from bs4 import Comment comments = stats.find_all(string=lambda text: isinstance(text, Comment)) for comment in comments: comment_soup = BeautifulSoup(comment, 'html.parser') table = comment_soup.find('table') if table: break if not table: return render_template('error.html', message='无法找到球员的Per Game数据表格') # 提取表头 table_headings = [] thead = table.find('thead') if thead: header_row = thead.find('tr') table_headings = [th.get_text(strip=True) for th in header_row.find_all('th')] else: header_row = table.find('tr') if header_row: table_headings = [th.get_text(strip=True) for th in header_row.find_all('th')] # 提取表格数据 table_data = [] tbody = table.find('tbody') rows = tbody.find_all('tr') if tbody else table.find_all('tr')[1:] for row in rows: # 跳过重复的表头行 if row.get('class') and 'thead' in row.get('class'): continue row_data = [td.get_text(strip=True) for td in row.find_all('td')] table_data.append(row_data) return render_template('stats.html', player_name=player_name, headings=table_headings, data=table_data) else: return render_template('index.html', players=players)
关键注意点
- 必须加请求头,避免被网站反爬机制拦截。
- Basketball Reference的表格常藏在HTML注释中,这是你之前找不到
table的主要原因。 - 提取表头优先用
thead标签,比直接找tr更准确。 - 全程添加异常判断,避免因页面结构变化导致程序崩溃。
内容的提问来源于stack exchange,提问作者SumeetS413
相关产品推荐
相关产品推荐

