爬取ICC板球排名时相同class的积分、场次字段如何用BeautifulSoup区分提取
拆分同class场次、积分数据的解决方案
核心逻辑:ICC排名页的表格结构固定,每一行内table-body__cell u-center-text类的两个<td>永远按「比赛场次、积分」的顺序排列,直接按索引取值即可拆分两类数据。
修正后可直接运行的完整代码
import requests from bs4 import BeautifulSoup url1 = "https://www.icc-cricket.com/rankings/mens/team-rankings/odi/" page = requests.get(url1) soup1 = BeautifulSoup(page.content, "html.parser") matches = [] points = [] # 处理排名第一的顶部横幅行 banner_matches = soup1.find("td", class_="rankings-block__banner-matches").text.strip() banner_points = soup1.find("td", class_="rankings-block__banner-points").text.strip() matches.append(banner_matches) points.append(banner_points) # 处理TOP10剩余9支队伍的普通表格行 for row in soup1.find_all("tr", class_="table-body")[:9]: # 提取当前行所有同class的居中单元格 center_cells = row.find_all("td", class_="table-body__cell u-center-text") # 按固定顺序拆分:索引0为场次,索引1为积分 matches.append(center_cells[0].text.strip()) points.append(center_cells[1].text.strip()) # 打印验证结果 print("TOP10队伍比赛场次:", matches) print("TOP10队伍对应积分:", points)
注意事项
- 你原有代码存在语法错误,
soup1 = BeautifulSoup(...)和print(soup1.prettify())不能写在同一行 - 原有代码仅提取了排名第一的队伍场次,没有处理剩余9支队伍的行,上述代码已补充该逻辑
- 限制
[:9]是为了只取TOP10数据,如果你需要爬取全榜数据可以去掉这个切片
内容的提问来源于stack exchange,提问作者bishwajit bhattacharya
相关产品推荐
相关产品推荐

