You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

爬取ICC板球排名时相同class的积分、场次字段如何用BeautifulSoup区分提取

拆分同class场次、积分数据的解决方案

核心逻辑:ICC排名页的表格结构固定,每一行内table-body__cell u-center-text类的两个<td>永远按「比赛场次、积分」的顺序排列,直接按索引取值即可拆分两类数据。

修正后可直接运行的完整代码

import requests
from bs4 import BeautifulSoup

url1 = "https://www.icc-cricket.com/rankings/mens/team-rankings/odi/"
page = requests.get(url1)
soup1 = BeautifulSoup(page.content, "html.parser")

matches = []
points = []

# 处理排名第一的顶部横幅行
banner_matches = soup1.find("td", class_="rankings-block__banner-matches").text.strip()
banner_points = soup1.find("td", class_="rankings-block__banner-points").text.strip()
matches.append(banner_matches)
points.append(banner_points)

# 处理TOP10剩余9支队伍的普通表格行
for row in soup1.find_all("tr", class_="table-body")[:9]:
    # 提取当前行所有同class的居中单元格
    center_cells = row.find_all("td", class_="table-body__cell u-center-text")
    # 按固定顺序拆分:索引0为场次,索引1为积分
    matches.append(center_cells[0].text.strip())
    points.append(center_cells[1].text.strip())

# 打印验证结果
print("TOP10队伍比赛场次:", matches)
print("TOP10队伍对应积分:", points)

注意事项

  • 你原有代码存在语法错误,soup1 = BeautifulSoup(...)和print(soup1.prettify())不能写在同一行
  • 原有代码仅提取了排名第一的队伍场次,没有处理剩余9支队伍的行,上述代码已补充该逻辑
  • 限制[:9]是为了只取TOP10数据,如果你需要爬取全榜数据可以去掉这个切片

内容的提问来源于stack exchange,提问作者bishwajit bhattacharya

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.02 08:15:04