You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python爬虫:遍历Target下拉列表并爬取对应页面表格

解决方案

你不需要模拟“点击”操作,直接通过修改URL中的target参数就能访问每个Target对应的页面,具体实现步骤如下:

  1. 构造每个Target的请求URL
    观察原URL结构:https://predictioncenter.org/casp14/results.cgi?view=tables&target=T1024&model=1&groups_id=,其中target=T1024是唯一需要替换的部分,把T1024换成你获取到的每个Target名称即可。

  2. 完整遍历爬取代码示例
    结合你已有的代码,扩展成完整的爬取逻辑:

import requests
from bs4 import BeautifulSoup
import time

# 基础URL模板,预留target参数位置
base_url = "https://predictioncenter.org/casp14/results.cgi?view=tables&target={}&model=1&groups_id="

# 请求头,模拟浏览器访问
headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36"
}

# 获取所有Target列表
init_response = requests.get(base_url.format("T1024"), headers=headers)
soup = BeautifulSoup(init_response.text, "html.parser")
options = soup.find("select", {"name": "target"}).findAll("option")
list_prot = [i.text for i in options]

# 遍历每个Target爬取表格
for target in list_prot:
    current_url = base_url.format(target)
    res = requests.get(current_url, headers=headers)
    res.encoding = "utf-8"
    current_soup = BeautifulSoup(res.text, "html.parser")
    
    # 定位目标表格(根据页面实际结构调整,这里以页面内第一个表格为例)
    target_table = current_soup.find("table")
    if target_table:
        print(f"=== 处理Target: {target} ===")
        # 提取表格行数据
        rows = target_table.find_all("tr")
        for row in rows:
            cols = row.find_all(["td", "th"])
            content = [col.get_text(strip=True) for col in cols]
            print(content)
        # 添加延迟,避免请求过于频繁
        time.sleep(1)
    else:
        print(f"Target {target} 未找到目标表格")
  1. 关键注意点
  • 必须带上User-Agent请求头,避免被网站拦截
  • 遍历过程中添加适当延迟,降低服务器访问压力
  • 解析表格时,要根据页面实际HTML结构调整find/find_all的参数,确保准确定位目标表格

内容的提问来源于stack exchange,提问作者ClaraJ

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.17 03:35:21