You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用BeautifulSoup无法提取ercot.com页面zip关联URL的问题

提取ERCOT页面中zip元素对应的URL问题

我需要提取Chrome右键审查元素可见的页面元素中的所有URL,目标页面地址:

url = fr'https://www.ercot.com/mp/data-products/data-product-details?id=NP6-788-CD'

目标URL是页面左侧“zip”选项审查元素时,右侧显示的URL(对应下图元素):
审查元素中的zip URL

我尝试了以下代码,但zip_urls1和zip_urls2均为空:

url = fr'https://www.ercot.com/mp/data-products/data-product-details?id=NP6-788-CD'
from bs4 import BeautifulSoup
from requests_html import HTMLSession
from shutil import copyfileobj
session = HTMLSession()
resp = session.get(url)
resp.html.render()
soup1 = BeautifulSoup(resp.html.html, "lxml").find_all("td")[::1]
zip_urls1 = [a.get('title') for a in soup1 if a.get('title') is not None]

soup = BeautifulSoup(resp.html.html, "lxml").find_all("a")
zip_urls2 = [a.get('href') for a in soup if 'doclookupId' in a.get('href')]

问题分析与修复

你的代码没拿到目标URL,核心是元素定位逻辑不对:

  • 目标zip链接不在<td>的title属性里,而是<td>内部<a>标签的href属性
  • 直接遍历所有<a>标签过滤doclookupId,可能因为页面渲染不充分或选择器太宽泛导致漏抓

修正后的代码

from bs4 import BeautifulSoup
from requests_html import HTMLSession

url = 'https://www.ercot.com/mp/data-products/data-product-details?id=NP6-788-CD'
session = HTMLSession()

# 渲染页面并等待2秒,确保动态内容完全加载
resp = session.get(url)
resp.html.render(sleep=2)

soup = BeautifulSoup(resp.html.html, "lxml")

# 精准定位包含zip文本的td,再提取内部的a标签链接
zip_urls = []
# 找所有文本包含zip的td标签(不区分大小写)
zip_tds = soup.find_all('td', string=lambda text: text and 'zip' in text.lower())
for td in zip_tds:
    link = td.find('a')
    if link and 'href' in link.attrs:
        zip_urls.append(link['href'])

# 打印结果
print(zip_urls)

补充说明

  • 新增sleep=2是给页面留足动态加载时间,避免元素还没生成就抓取
  • 通过文本匹配定位目标<td>,比盲目遍历更精准
  • 如果提取到的是相对路径,可以拼接成完整URL:full_url = 'https://www.ercot.com' + link['href']

内容的提问来源于stack exchange,提问作者Zanam

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.21 20:22:37