如何修复新手网页爬虫代码中的AttributeError错误?
修复NFL球员数据爬虫的AttributeError错误
问题原因
报错AttributeError: 'NoneType' object has no attribute 'tbody'是因为soup.find('table', {'id': 'players'})返回了None(找不到目标表格),同时代码存在严重缩进错误,导致循环逻辑完全混乱,无法正确遍历多年份页面并提取数据。
具体修复步骤
- 修正缩进逻辑:将遍历行的
for row循环、数据提取代码全部缩进至年份循环内部,确保每个年份的页面都被完整处理。 - 增加请求有效性检查:解析页面之前先判断请求是否成功(
response.status_code == 200),避免解析错误页面。 - 增加表格查找容错:确认找到目标表格后再访问其
tbody属性,避免直接调用None的方法。 - 清理冗余代码:移除未定义的数据库操作代码(
c.execute、getAll、querySave),若需要数据库操作,请提前初始化连接和相关函数。
修复后的完整代码
import requests from bs4 import BeautifulSoup # Base URL for the NFL player stats page base_url = 'https://www.pro-football-reference.com/players/' # List to store player data player_data = [] # Loop through the years 2014 to 2021 for year in range(2014, 2022): # Send a GET request to the URL url = f'{base_url}{year}/' response = requests.get(url) # 检查请求是否成功 if response.status_code != 200: print(f"请求{url}失败,状态码:{response.status_code}") continue # Parse the HTML of the page soup = BeautifulSoup(response.text, 'html.parser') # Find the player stats table, 增加容错判断 player_table = soup.find('table', {'id': 'players'}) if not player_table: print(f"在{url}中未找到id为players的表格") continue # Find all rows in the player stats table rows = player_table.tbody.find_all('tr')[1:] # Loop through each row(缩进至年份循环内) for row in rows: # Find the player name cell name_cell = row.find('th') # Check if the cell is valid (some rows may not have player data) if name_cell: # Extract the player name and link try: name = name_cell.a.text except AttributeError: name = '' try: position = row.find('td', {'data-stat': 'position'}).text except AttributeError: position = '' try: link = name_cell.a['href'] except AttributeError: link = '' # Extract the player height and weight try: height = row.find('td', {'data-stat': 'height'}).text except AttributeError: height = '' try: weight = row.find('td', {'data-stat': 'weight'}).text except AttributeError: weight = '' # Add the player data to the list(缩进至if name_cell内) player_data.append({ 'name': name, 'position': position, 'link': link, 'height': height, 'weight': weight }) # Print the player data print(player_data) # 若需要数据库操作,请先初始化数据库连接和相关函数 # c.execute(...) # getAll('player_data',c) # querySave(player_data, c, 'NFLHeightWeight') print("Done!")
额外说明
- 如果仍出现找不到表格的情况,可能是网站启用了反爬机制,可尝试添加请求头模拟浏览器请求:
headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36' } response = requests.get(url, headers=headers) - 建议添加延迟(比如
time.sleep(1))避免频繁请求触发反爬。
内容的提问来源于stack exchange,提问作者Josh Scott
相关产品推荐
相关产品推荐

