You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python爬虫问题:如何定位Liquipedia无类名div的选手信息

爬取Liquipedia s1mple选手信息模块

Liquipedia的选手页面中,"Player Information"模块的结构有固定规律——标题对应带id="Player_Information"的<span>标签,后续紧跟的表格就是存储键值对的内容容器。以下是直接的爬取和整理方案:

代码实现

import requests
from bs4 import BeautifulSoup

# 目标页面URL
url = "https://liquipedia.net/counterstrike/s1mple"

# 模拟浏览器请求头,规避基础反爬
headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36"
}

# 获取并解析页面
response = requests.get(url, headers=headers)
response.encoding = "utf-8"
soup = BeautifulSoup(response.text, "html.parser")

# 定位模块标题(通过锚点ID是最稳定的方式)
info_title = soup.find("span", id="Player_Information")
if not info_title:
    print("未找到Player Information模块")
    exit()

# 获取标题所在h2标签的后续表格(模块内容容器)
info_table = info_title.parent.find_next_sibling("table")

# 提取键值对,整理为两列结构
player_info = []
for row in info_table.find_all("tr"):
    cells = row.find_all("td")
    if len(cells) == 2:
        label = cells[0].get_text(strip=True)
        value = cells[1].get_text(strip=True)
        player_info.append([label, value])

# 打印两列格式的结果
print("Player Information(两列结构):")
for item in player_info:
    print(f"{item[0]:<20} | {item[1]}")

关键说明

  • 用id="Player_Information"定位标题比纯文本匹配更可靠,页面锚点ID不会轻易变动
  • 过滤单元格数量为2的行,避免表头、空行干扰
  • get_text(strip=True)自动去除文本前后的空格、换行,让内容更整洁
  • 若页面结构后续调整(比如表格换成<dl>定义列表),可修改逻辑为:找到标题后,遍历后续<dl>中的<dt>(键)和<dd>(值)

内容的提问来源于stack exchange,提问作者quarantinho

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.30 08:35:01