You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python爬取足球球员数据:如何获取位置、市值等字段?

爬取Ligainsider球员完整数据的解决方案

问题概述

我正在为个人项目爬取https://www.ligainsider.de/stats/kickbase/rangliste/feldspieler/gesamt/上的足球球员数据表,目标是抓取页面完整表格数据。目前已通过Python的BeautifulSoup成功获取球员姓名(Spieler)、球队(Verein)、单轮得分(Punkte Spieltag)数据,但因多个字段使用相同的text-right类标签,无法准确定位位置(Position)、市值(Marktwert)、出场次数(Einsätze)等字段。

现有代码

from bs4 import BeautifulSoup
import requests
import pandas as pd

# Define configuration variables
URL = "https://www.ligainsider.de/stats/kickbase/rangliste/feldspieler/gesamt/"
output_path = "C:[...].xlsx"

# Make request to the website
page = requests.get(URL)

# Parse the HTML content using BeautifulSoup
soup = BeautifulSoup(page.content, "html.parser")

# Find all rows with information
table_rows = soup.find_all("tr")

# Create empty lists to hold the data
player_names = []
team_names = []
points_matchday = []

# Loop through each row and extract the data
for row in table_rows:
    # Extract player name
    player_name = row.find("a")
    if player_name is not None:
        player_name_text = player_name.text.strip()
        player_names.append(player_name_text)

    # Extract team name
    team_name = row.find("a", {"class", "text-thin"})
    if team_name is not None:
        team_name_text = team_name.text.strip()
        team_names.append(team_name_text)

    # Extract last matchday points
    points_matchday_name = row.find("td", {"class", "text-right"})
    if points_matchday_name is not None:
        points_matchday_name_text = points_matchday_name.text.strip()
        points_matchday.append(points_matchday_name_text)

print(player_names)
print(team_names)
print(points_matchday)

目标行HTML示例(以Joshua Kimmich为例)

<tr data-anchor-rowfilter="filter1" role="row" class="odd">

    <td data-criteria-rowfilter="" style="display: none;">
        <ul>
            <li>joshuakimmich</li>
            <li>joshuakimmich</li>
            <li>fcbayernmunchen</li>
            <li>fcbayernmünchen</li>
            <li>mittelfeldspieler</li>
        </ul>
    </td>
    <td class="sorting_1">1</td>
    <td class="text-left">
        <strong><a href="/joshua-kimmich_5768/">Joshua Kimmich</a></strong>
    </td>
    <td class="text-left"><a class="text-thin" href="/fc-bayern-muenchen/1/">FC Bayern München</a></td>
    <td class="text-left">Mittelfeldspieler</td>
        <td class="text-right">164</td>
    <td class="text-right">50.432.349€</td>
    <td class="text-right">21</td>
    <td class="text-right">162,10</td>
    <td class="text-right"><strong>3.404</strong></td>
    
</tr>

解决方案

问题出在使用find()方法获取text-right标签时,只会返回第一个匹配项,而表格中有多个text-right类的单元格。正确的做法是获取当前行的所有<td>元素,然后根据固定索引提取对应字段(表格结构是固定的)。

修改后的代码如下:

from bs4 import BeautifulSoup
import requests
import pandas as pd

URL = "https://www.ligainsider.de/stats/kickbase/rangliste/feldspieler/gesamt/"
output_path = "C:[...].xlsx"

page = requests.get(URL)
soup = BeautifulSoup(page.content, "html.parser")
table_rows = soup.find_all("tr")

# 初始化所有需要的字段列表
player_names = []
team_names = []
positions = []
total_points = []
market_values = []
appearances = []
avg_points = []
matchday_points = []

for row in table_rows:
    # 获取当前行所有td元素
    tds = row.find_all("td")
    # 跳过表头和空行(有效数据行至少有10个td)
    if len(tds) < 10:
        continue
    
    # 按索引提取对应字段
    # 球员姓名:tds[2]中的a标签文本
    player_name = tds[2].find("a").text.strip()
    player_names.append(player_name)
    
    # 球队名称:tds[3]中的a标签文本
    team_name = tds[3].find("a").text.strip()
    team_names.append(team_name)
    
    # 位置:tds[4]的文本
    position = tds[4].text.strip()
    positions.append(position)
    
    # 总得分:tds[5]的文本
    total_point = tds[5].text.strip()
    total_points.append(total_point)
    
    # 市值:tds[6]的文本
    market_value = tds[6].text.strip()
    market_values.append(market_value)
    
    # 出场次数:tds[7]的文本
    appearance = tds[7].text.strip()
    appearances.append(appearance)
    
    # 场均得分:tds[8]的文本
    avg_point = tds[8].text.strip()
    avg_points.append(avg_point)
    
    # 单轮得分:tds[9]的文本
    matchday_point = tds[9].text.strip()
    matchday_points.append(matchday_point)

# 整理成DataFrame并保存到Excel
df = pd.DataFrame({
    "球员姓名": player_names,
    "球队": team_names,
    "位置": positions,
    "总得分": total_points,
    "市值": market_values,
    "出场次数": appearances,
    "场均得分": avg_points,
    "单轮得分": matchday_points
})

df.to_excel(output_path, index=False)
print("数据已成功保存到", output_path)

关键说明

  • 通过row.find_all("td")获取当前行所有单元格,确保能按顺序访问每个字段
  • 跳过长度不足10的行(表头和空行),避免索引越界
  • 每个字段对应固定的td索引(依据提供的Kimmich行HTML结构确定):
    • 索引2:球员姓名
    • 索引3:球队
    • 索引4:位置
    • 索引5:总得分
    • 索引6:市值
    • 索引7:出场次数
    • 索引8:场均得分
    • 索引9:单轮得分

内容的提问来源于stack exchange,提问作者Marco

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.29 19:12:08