网页抓取为何出现NaN值?NHL球员数据爬取代码求助
问题分析与解决方案
核心问题点
- 代码中
name变量未定义,导致无法匹配球员姓名填充链接,最终返回NaN - 链接遍历逻辑错误,
list变量被重复赋值为单个链接,无法存储所有链接 - 未在遍历表格行的同时关联球员链接,导致姓名与链接无法一一对应
修正后的代码
import requests from bs4 import BeautifulSoup import pandas as pd # 目标赛季URL start_url = 'https://www.eliteprospects.com/league/nhl/stats/2023-2024' r = requests.get(url=start_url) soup = BeautifulSoup(r.content, "html.parser") table = soup.find("table", class_="table table-striped table-sortable player-stats highlight-stats season") # 提取表头 headers = table.find_all("th") titles = [i.text.strip() for i in headers] # 添加球员链接列 titles.append("Player_URL") df = pd.DataFrame(columns=titles) rows = table.find_all("tr")[1:] # 跳过表头行 for row in rows: # 提取当前行的所有单元格数据 cells = row.find_all("td") row_data = [cell.text.strip().replace('\n', '') for cell in cells] # 提取球员个人页面链接 player_link_elem = row.find("span", class_="txt-blue").find_parent("a") if player_link_elem: player_url = f"https://www.eliteprospects.com{player_link_elem.get('href')}" row_data.append(player_url) # 将数据添加到DataFrame df.loc[len(df)] = row_data # 设置显示参数 pd.set_option('display.max_columns', None) pd.set_option('display.max_rows', None) pd.set_option('display.width', 999) pd.set_option('display.expand_frame_repr', True) print(df.to_string())
关键修正说明
- 遍历表格行时,直接从当前行查找球员链接元素,确保姓名与链接一一对应
- 补全链接的完整URL(原页面中是相对路径)
- 提前添加
Player_URL列到表头,避免后续列不匹配 - 统一处理文本的换行和空格,保证数据整洁
内容的提问来源于stack exchange,提问作者ijif
相关产品推荐
相关产品推荐

