如何通过Python获取JavaScript页面中DATA V CSV对应的CSV链接?
解决无法获取「DATA V CSV」对应CSV链接的问题
可能的原因
- 元素文本匹配不精准:
string='DATA V CSV'是完全匹配,若文本有多余空格、换行或被嵌套标签包裹,会匹配失败 - 链接未直接存储在
href属性:按钮可能通过onclick事件触发下载,链接藏在JavaScript代码中 - 动态加载内容:页面部分内容依赖JavaScript渲染,
requests获取的静态HTML中不存在目标元素
解决方案
方案1:改进静态爬取逻辑
通过模糊匹配文本、限定目标区域、解析事件属性来获取链接:
import requests import re from bs4 import BeautifulSoup url = 'https://www.ceps.cz/en/all-data#AktualniSystemovaOdchylkaCR' response = requests.get(url) soup = BeautifulSoup(response.content, 'html.parser') # 定位到锚点对应的目标区域 target_section = soup.find(id='AktualniSystemovaOdchylkaCR') if target_section: # 在目标区域内查找包含目标文本的按钮或链接 csv_element = target_section.find( lambda tag: tag.name in ['a', 'button'] and 'DATA V CSV' in tag.get_text(strip=True) ) if csv_element: # 优先获取href属性 if 'href' in csv_element.attrs: csv_link = csv_element['href'] print(f"找到CSV链接:{csv_link}") # 解析onclick事件中的链接 elif 'onclick' in csv_element.attrs: onclick_content = csv_element['onclick'] link_match = re.search(r"'(https?://[^']+)'", onclick_content) if link_match: csv_link = link_match.group(1) print(f"找到CSV链接:{csv_link}") else: print("未从onclick事件中提取到链接") else: print("目标元素无有效链接属性") else: print("未找到包含「DATA V CSV」文本的元素") else: print("未找到目标锚点对应的区域")
方案2:使用Selenium处理动态内容
若页面依赖JavaScript渲染目标元素,需模拟浏览器加载完整页面:
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC import re url = 'https://www.ceps.cz/en/all-data#AktualniSystemovaOdchylkaCR' # 初始化浏览器驱动(需提前安装对应浏览器的驱动并配置环境变量) driver = webdriver.Chrome() driver.get(url) try: # 等待目标元素加载完成,最多等待10秒 csv_button = WebDriverWait(driver, 10).until( EC.presence_of_element_located((By.XPATH, "//*[contains(text(), 'DATA V CSV')]")) ) # 获取链接 csv_link = None if csv_button.get_attribute('href'): csv_link = csv_button.get_attribute('href') else: onclick_content = csv_button.get_attribute('onclick') link_match = re.search(r"'(https?://[^']+)'", onclick_content) if link_match: csv_link = link_match.group(1) if csv_link: print(f"找到CSV链接:{csv_link}") else: print("未找到有效链接") finally: # 关闭浏览器 driver.quit()
内容的提问来源于stack exchange,提问作者user189035
相关产品推荐
相关产品推荐

