You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何通过Python获取JavaScript页面中DATA V CSV对应的CSV链接?

解决无法获取「DATA V CSV」对应CSV链接的问题

可能的原因

  • 元素文本匹配不精准:string='DATA V CSV'是完全匹配,若文本有多余空格、换行或被嵌套标签包裹,会匹配失败
  • 链接未直接存储在href属性:按钮可能通过onclick事件触发下载,链接藏在JavaScript代码中
  • 动态加载内容:页面部分内容依赖JavaScript渲染,requests获取的静态HTML中不存在目标元素

解决方案

方案1:改进静态爬取逻辑

通过模糊匹配文本、限定目标区域、解析事件属性来获取链接:

import requests
import re
from bs4 import BeautifulSoup

url = 'https://www.ceps.cz/en/all-data#AktualniSystemovaOdchylkaCR'
response = requests.get(url)

soup = BeautifulSoup(response.content, 'html.parser')

# 定位到锚点对应的目标区域
target_section = soup.find(id='AktualniSystemovaOdchylkaCR')
if target_section:
    # 在目标区域内查找包含目标文本的按钮或链接
    csv_element = target_section.find(
        lambda tag: tag.name in ['a', 'button'] and 'DATA V CSV' in tag.get_text(strip=True)
    )
    if csv_element:
        # 优先获取href属性
        if 'href' in csv_element.attrs:
            csv_link = csv_element['href']
            print(f"找到CSV链接:{csv_link}")
        # 解析onclick事件中的链接
        elif 'onclick' in csv_element.attrs:
            onclick_content = csv_element['onclick']
            link_match = re.search(r"'(https?://[^']+)'", onclick_content)
            if link_match:
                csv_link = link_match.group(1)
                print(f"找到CSV链接:{csv_link}")
            else:
                print("未从onclick事件中提取到链接")
        else:
            print("目标元素无有效链接属性")
    else:
        print("未找到包含「DATA V CSV」文本的元素")
else:
    print("未找到目标锚点对应的区域")

方案2:使用Selenium处理动态内容

若页面依赖JavaScript渲染目标元素,需模拟浏览器加载完整页面:

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
import re

url = 'https://www.ceps.cz/en/all-data#AktualniSystemovaOdchylkaCR'

# 初始化浏览器驱动(需提前安装对应浏览器的驱动并配置环境变量)
driver = webdriver.Chrome()
driver.get(url)

try:
    # 等待目标元素加载完成,最多等待10秒
    csv_button = WebDriverWait(driver, 10).until(
        EC.presence_of_element_located((By.XPATH, "//*[contains(text(), 'DATA V CSV')]"))
    )
    # 获取链接
    csv_link = None
    if csv_button.get_attribute('href'):
        csv_link = csv_button.get_attribute('href')
    else:
        onclick_content = csv_button.get_attribute('onclick')
        link_match = re.search(r"'(https?://[^']+)'", onclick_content)
        if link_match:
            csv_link = link_match.group(1)
    if csv_link:
        print(f"找到CSV链接:{csv_link}")
    else:
        print("未找到有效链接")
finally:
    # 关闭浏览器
    driver.quit()

内容的提问来源于stack exchange,提问作者user189035

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.26 23:12:22