Python爬虫:遍历Target下拉列表并爬取对应页面表格
解决方案
你不需要模拟“点击”操作,直接通过修改URL中的target参数就能访问每个Target对应的页面,具体实现步骤如下:
构造每个Target的请求URL
观察原URL结构:https://predictioncenter.org/casp14/results.cgi?view=tables&target=T1024&model=1&groups_id=,其中target=T1024是唯一需要替换的部分,把T1024换成你获取到的每个Target名称即可。完整遍历爬取代码示例
结合你已有的代码,扩展成完整的爬取逻辑:
import requests from bs4 import BeautifulSoup import time # 基础URL模板,预留target参数位置 base_url = "https://predictioncenter.org/casp14/results.cgi?view=tables&target={}&model=1&groups_id=" # 请求头,模拟浏览器访问 headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36" } # 获取所有Target列表 init_response = requests.get(base_url.format("T1024"), headers=headers) soup = BeautifulSoup(init_response.text, "html.parser") options = soup.find("select", {"name": "target"}).findAll("option") list_prot = [i.text for i in options] # 遍历每个Target爬取表格 for target in list_prot: current_url = base_url.format(target) res = requests.get(current_url, headers=headers) res.encoding = "utf-8" current_soup = BeautifulSoup(res.text, "html.parser") # 定位目标表格(根据页面实际结构调整,这里以页面内第一个表格为例) target_table = current_soup.find("table") if target_table: print(f"=== 处理Target: {target} ===") # 提取表格行数据 rows = target_table.find_all("tr") for row in rows: cols = row.find_all(["td", "th"]) content = [col.get_text(strip=True) for col in cols] print(content) # 添加延迟,避免请求过于频繁 time.sleep(1) else: print(f"Target {target} 未找到目标表格")
- 关键注意点
- 必须带上
User-Agent请求头,避免被网站拦截 - 遍历过程中添加适当延迟,降低服务器访问压力
- 解析表格时,要根据页面实际HTML结构调整
find/find_all的参数,确保准确定位目标表格
内容的提问来源于stack exchange,提问作者ClaraJ
相关产品推荐
相关产品推荐

