You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python提取指定网页的URL及Number等特定列数据?

解决方案:同步提取表格数据与下载链接

实现思路

直接用BeautifulSoup解析网页表格,同时提取每一行的列数据和Network列对应的下载链接,最后整理成Pandas DataFrame,兼顾数据完整性和链接获取需求。

代码示例

import requests
from bs4 import BeautifulSoup
import pandas as pd

# 目标网页地址
url = "https://training.lczero.org/networks/?show_all=1"

# 获取网页内容
response = requests.get(url)
response.raise_for_status()  # 捕获请求失败的异常

# 解析HTML结构
soup = BeautifulSoup(response.text, "html.parser")

# 定位页面中的目标表格
table = soup.find("table")

# 提取表头文本并筛选目标列的索引
headers = [th.get_text(strip=True) for th in table.find("thead").find_all("th")]
target_cols = ["Number", "Run", "Network", "Elo", "Games"]
col_indices = [headers.index(col) for col in target_cols]

# 初始化存储数据的列表
data_rows = []

# 遍历表格所有行(跳过表头行)
for row in table.find("tbody").find_all("tr"):
    cells = row.find_all("td")
    # 提取目标列的文本内容
    row_content = [cells[idx].get_text(strip=True) for idx in col_indices]
    # 从Network列的a标签中提取下载链接
    network_link = cells[headers.index("Network")].find("a")["href"]
    # 将链接追加到当前行数据中
    row_content.append(network_link)
    data_rows.append(row_content)

# 构造DataFrame,新增一列存储下载链接
result_df = pd.DataFrame(data_rows, columns=target_cols + ["Download Link"])

# 打印前5行验证结果
print(result_df.head())

代码说明

  • 先通过requests拉取网页源码,确保请求成功后再用BeautifulSoup解析
  • 定位页面中的表格,提取表头并匹配需要的列对应的位置索引
  • 逐行提取目标列的文本内容,同时从Network列的<a>标签中抓取下载链接
  • 将所有数据整理成DataFrame,新增Download Link列单独存储链接,方便后续批量下载或数据处理

内容的提问来源于stack exchange,提问作者Meet Shah

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.15 23:45:40