Python爬取足球球员数据:如何获取位置、市值等字段?
爬取Ligainsider球员完整数据的解决方案
问题概述
我正在为个人项目爬取https://www.ligainsider.de/stats/kickbase/rangliste/feldspieler/gesamt/上的足球球员数据表,目标是抓取页面完整表格数据。目前已通过Python的BeautifulSoup成功获取球员姓名(Spieler)、球队(Verein)、单轮得分(Punkte Spieltag)数据,但因多个字段使用相同的text-right类标签,无法准确定位位置(Position)、市值(Marktwert)、出场次数(Einsätze)等字段。
现有代码
from bs4 import BeautifulSoup import requests import pandas as pd # Define configuration variables URL = "https://www.ligainsider.de/stats/kickbase/rangliste/feldspieler/gesamt/" output_path = "C:[...].xlsx" # Make request to the website page = requests.get(URL) # Parse the HTML content using BeautifulSoup soup = BeautifulSoup(page.content, "html.parser") # Find all rows with information table_rows = soup.find_all("tr") # Create empty lists to hold the data player_names = [] team_names = [] points_matchday = [] # Loop through each row and extract the data for row in table_rows: # Extract player name player_name = row.find("a") if player_name is not None: player_name_text = player_name.text.strip() player_names.append(player_name_text) # Extract team name team_name = row.find("a", {"class", "text-thin"}) if team_name is not None: team_name_text = team_name.text.strip() team_names.append(team_name_text) # Extract last matchday points points_matchday_name = row.find("td", {"class", "text-right"}) if points_matchday_name is not None: points_matchday_name_text = points_matchday_name.text.strip() points_matchday.append(points_matchday_name_text) print(player_names) print(team_names) print(points_matchday)
目标行HTML示例(以Joshua Kimmich为例)
<tr data-anchor-rowfilter="filter1" role="row" class="odd"> <td data-criteria-rowfilter="" style="display: none;"> <ul> <li>joshuakimmich</li> <li>joshuakimmich</li> <li>fcbayernmunchen</li> <li>fcbayernmünchen</li> <li>mittelfeldspieler</li> </ul> </td> <td class="sorting_1">1</td> <td class="text-left"> <strong><a href="/joshua-kimmich_5768/">Joshua Kimmich</a></strong> </td> <td class="text-left"><a class="text-thin" href="/fc-bayern-muenchen/1/">FC Bayern München</a></td> <td class="text-left">Mittelfeldspieler</td> <td class="text-right">164</td> <td class="text-right">50.432.349€</td> <td class="text-right">21</td> <td class="text-right">162,10</td> <td class="text-right"><strong>3.404</strong></td> </tr>
解决方案
问题出在使用find()方法获取text-right标签时,只会返回第一个匹配项,而表格中有多个text-right类的单元格。正确的做法是获取当前行的所有<td>元素,然后根据固定索引提取对应字段(表格结构是固定的)。
修改后的代码如下:
from bs4 import BeautifulSoup import requests import pandas as pd URL = "https://www.ligainsider.de/stats/kickbase/rangliste/feldspieler/gesamt/" output_path = "C:[...].xlsx" page = requests.get(URL) soup = BeautifulSoup(page.content, "html.parser") table_rows = soup.find_all("tr") # 初始化所有需要的字段列表 player_names = [] team_names = [] positions = [] total_points = [] market_values = [] appearances = [] avg_points = [] matchday_points = [] for row in table_rows: # 获取当前行所有td元素 tds = row.find_all("td") # 跳过表头和空行(有效数据行至少有10个td) if len(tds) < 10: continue # 按索引提取对应字段 # 球员姓名:tds[2]中的a标签文本 player_name = tds[2].find("a").text.strip() player_names.append(player_name) # 球队名称:tds[3]中的a标签文本 team_name = tds[3].find("a").text.strip() team_names.append(team_name) # 位置:tds[4]的文本 position = tds[4].text.strip() positions.append(position) # 总得分:tds[5]的文本 total_point = tds[5].text.strip() total_points.append(total_point) # 市值:tds[6]的文本 market_value = tds[6].text.strip() market_values.append(market_value) # 出场次数:tds[7]的文本 appearance = tds[7].text.strip() appearances.append(appearance) # 场均得分:tds[8]的文本 avg_point = tds[8].text.strip() avg_points.append(avg_point) # 单轮得分:tds[9]的文本 matchday_point = tds[9].text.strip() matchday_points.append(matchday_point) # 整理成DataFrame并保存到Excel df = pd.DataFrame({ "球员姓名": player_names, "球队": team_names, "位置": positions, "总得分": total_points, "市值": market_values, "出场次数": appearances, "场均得分": avg_points, "单轮得分": matchday_points }) df.to_excel(output_path, index=False) print("数据已成功保存到", output_path)
关键说明
- 通过
row.find_all("td")获取当前行所有单元格,确保能按顺序访问每个字段 - 跳过长度不足10的行(表头和空行),避免索引越界
- 每个字段对应固定的td索引(依据提供的Kimmich行HTML结构确定):
- 索引2:球员姓名
- 索引3:球队
- 索引4:位置
- 索引5:总得分
- 索引6:市值
- 索引7:出场次数
- 索引8:场均得分
- 索引9:单轮得分
内容的提问来源于stack exchange,提问作者Marco
相关产品推荐
相关产品推荐

