使用BeautifulSoup无法提取ercot.com页面zip关联URL的问题
提取ERCOT页面中zip元素对应的URL问题
我需要提取Chrome右键审查元素可见的页面元素中的所有URL,目标页面地址:
url = fr'https://www.ercot.com/mp/data-products/data-product-details?id=NP6-788-CD'
目标URL是页面左侧“zip”选项审查元素时,右侧显示的URL(对应下图元素):
我尝试了以下代码,但zip_urls1和zip_urls2均为空:
url = fr'https://www.ercot.com/mp/data-products/data-product-details?id=NP6-788-CD' from bs4 import BeautifulSoup from requests_html import HTMLSession from shutil import copyfileobj session = HTMLSession() resp = session.get(url) resp.html.render() soup1 = BeautifulSoup(resp.html.html, "lxml").find_all("td")[::1] zip_urls1 = [a.get('title') for a in soup1 if a.get('title') is not None] soup = BeautifulSoup(resp.html.html, "lxml").find_all("a") zip_urls2 = [a.get('href') for a in soup if 'doclookupId' in a.get('href')]
问题分析与修复
你的代码没拿到目标URL,核心是元素定位逻辑不对:
- 目标zip链接不在
<td>的title属性里,而是<td>内部<a>标签的href属性 - 直接遍历所有
<a>标签过滤doclookupId,可能因为页面渲染不充分或选择器太宽泛导致漏抓
修正后的代码
from bs4 import BeautifulSoup from requests_html import HTMLSession url = 'https://www.ercot.com/mp/data-products/data-product-details?id=NP6-788-CD' session = HTMLSession() # 渲染页面并等待2秒,确保动态内容完全加载 resp = session.get(url) resp.html.render(sleep=2) soup = BeautifulSoup(resp.html.html, "lxml") # 精准定位包含zip文本的td,再提取内部的a标签链接 zip_urls = [] # 找所有文本包含zip的td标签(不区分大小写) zip_tds = soup.find_all('td', string=lambda text: text and 'zip' in text.lower()) for td in zip_tds: link = td.find('a') if link and 'href' in link.attrs: zip_urls.append(link['href']) # 打印结果 print(zip_urls)
补充说明
- 新增
sleep=2是给页面留足动态加载时间,避免元素还没生成就抓取 - 通过文本匹配定位目标
<td>,比盲目遍历更精准 - 如果提取到的是相对路径,可以拼接成完整URL:
full_url = 'https://www.ercot.com' + link['href']
内容的提问来源于stack exchange,提问作者Zanam
相关产品推荐
相关产品推荐

