如何用Beautiful Soup爬取NHL阵容并生成指定格式DataFrame?
解决NHL阵容爬取并生成DataFrame的问题
要把球队、阵容线和球员、位置对应起来,得顺着页面的结构层级遍历——先定位每个球队的区块,再在球队内找每条阵容线,最后提取线内的位置和球员信息。修改后的完整代码如下:
import requests import pandas as pd from bs4 import BeautifulSoup url = "https://www.rotowire.com/hockey/nhl-lineups.php" response = requests.get(url) soup = BeautifulSoup(response.text, "html.parser") # 存储所有阵容数据的列表 lineup_data = [] # 获取所有球队的阵容区块 team_blocks = soup.find_all('div', class_='lineup__team') for team_block in team_blocks: # 提取球队名称 team_name = team_block.find('div', class_='lineup__team-name').text.strip() # 获取当前球队的所有阵容线(前锋线、后卫线、门将) line_blocks = team_block.find_all('div', class_='lineup__line') for line_block in line_blocks: # 提取阵容线编号(比如"Line 1"、"D Pair 1"、"Goalies") line_label = line_block.find('div', class_='lineup__line-label').text.strip() # 获取这条线里的所有位置-球员组合 pos_player_pairs = line_block.find_all('div', class_='lineup__pos-player') for pair in pos_player_pairs: # 提取位置 position = pair.find('div', class_='lineup__pos').text.strip() # 提取球员姓名(从title属性获取,和你原来的逻辑一致) player = pair.find('a', title=True)['title'] # 将数据添加到列表 lineup_data.append({ 'Team': team_name, 'Position': position, 'Player': player, 'Line': line_label }) # 转换为DataFrame df = pd.DataFrame(lineup_data) print(df.head())
关键逻辑说明:
- 层级遍历:从球队区块到单条阵容线,再到每个位置-球员对,确保每个球员都能关联到所属的球队和阵容线。
- 数据收集:用字典列表存储每条记录,最后直接转成DataFrame,保证数据结构一致。
- 标签定位:利用页面的class属性精准定位元素,比如
lineup__team对应球队、lineup__line-label对应阵容线名称,避免提取无关数据。
内容的提问来源于stack exchange,提问作者user1389739
相关产品推荐
相关产品推荐

