You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何爬取含on-click按钮网站的PDF链接?BeautifulSoup/lxml遇阻求助

问题描述

尝试用BeautifulSoup和lxml爬取http://www.italgiure.giustizia.it/sncass/的PDF文件,无法获取<span class="toDocument pdf">标签data-arg属性中的PDF链接,代码返回空列表。相关代码、运行结果及渲染后的HTML片段如下:

原代码

import requests
from bs4 import BeautifulSoup
from lxml import html

r= requests.get('http://www.italgiure.giustizia.it/sncass/')
soup = BeautifulSoup(r.text, 'html.parser')

pdf_list = soup.find_all('a')
print(pdf_list)
search_html = html.fromstring(r.text)
page_link = search_html.xpath('//*[@id="contentData"]/div[2]/div[1]/div/h3/a/span[1]/span')
print(page_link)

运行结果

[<a href="accessibilita.html" style="text-decoration:none;font-size:80%;color:white" tabindex="0">Accessibilità</a>, <a accesskey="r" name="results" onclick="$(this).next().focus();" tabindex="-2" title="contenuto"></a>, <a accesskey="1" name="card" onclick="$(this).next().focus();" tabindex="-2" title="documento"></a>, <a class="text2pdf" href="javascript:void(0)" onclick="toTargetDoc($('.toDocument.pdf',$(this)).attr('data-arg'), this)" style="text-decoration:none;color:#440;" tabindex="0"> <span data-arg="filename" data-role="content" title="pdf"></span>  <span class="chkcontent"><span class="label">Sez.</span> <span class="risultato" data-arg="szdec" data-role="content"></span> <span class="risultato" data-arg="kind" data-role="content"></span><span class="chkcontent"> - <span class="risultato" data-arg="ssz" data-role="content"></span></span><span class="label">,</span> </span> <span data-arg="tipoprov" data-role="content"></span> <span class="chkcontent"><span class="chkcontent"><span class="label">n.</span><span data-arg="numcard" data-role="content"></span></span><span data-arg="numdec" data-role="content" style="display:none"></span><span data-arg="numdep" data-role="content" style="display:none"></span> <span class="chkcontent"><span class="label"> del </span><span data-arg="datdep" data-role="content"></span><span data-arg="ecli" data-role="content" style="font-weight:normal"></span><span data-arg="anno" data-role="content" style="display:none"></span><span class="label">,</span></span> </span> <span class="chkcontent"><span class="label">udienza del</span> <span data-arg="datdec" data-role="content"></span><span class="label">,</span></span> <span class="chkcontent"><span class="label">Presidente </span><span data-arg="presidente" data-role="content"></span> </span> <span class="chkcontent"><span class="label">Relatore </span><span data-arg="relatore" data-role="content"></span> </span> </a>, <a class="text2ocr" href="javascript:void(0)" onclick="toTargetText($('.toDocument.txt',$(this)).attr('data-arg'))" style="text-decoration:none;color:#440;" tabindex="0"> <span data-arg="testoocr" data-role="content" title="testo ocr"></span>  <span data-arg="ocr" data-role="datasubset"> <span data-arg="ocr" data-role="multivaluedcontent">snippet</span> </span> </a>, <a href="http://www.italgiure.giustizia.it" style="color:white;" tabindex="0">ItalgiureWeb</a>] []

渲染后的目标HTML片段

<a href="javascript:void(0)" tabindex="0" onclick="toTargetDoc($('.toDocument.pdf',$(this)).attr('data-arg'), this)" style="text-decoration:none;color:#440;" class="text2pdf"> 
    <span data-role="content" data-arg="filename" title="pdf">
        <span class="toDocument pdf" data-arg="/xway/application/nif/clean/hc.dll%3Fverbo%3Dattach%26db%3Dsnciv%26id%3D./20221107/snciv@s50@a2022@n32765@tO.clean.pdf">
            <img class="rowIcon" alt="formato pdf" src="pix/pdf.png">
        </span>
    </span>&nbsp; 
    <span class="chkcontent"><span class="label">Sez.</span>&nbsp;<span class="risultato" data-role="content" data-arg="szdec">QUINTA</span> <span class="risultato" data-role="content" data-arg="kind">CIVILE</span><span class="label">,</span> </span> 
    <span data-role="content" data-arg="tipoprov">Ordinanza</span> 
    <span class="chkcontent"><span class="chkcontent"><span class="label">n.</span><span data-role="content" data-arg="numcard">32765</span></span><span style="display:none" data-role="content" data-arg="numdec">32765</span><span style="display:none" data-role="content" data-arg="numdep"></span> <span class="chkcontent"><span class="label"> del </span><span data-role="content" data-arg="datdep">07/11/2022</span><span style="font-weight:normal" data-role="content" data-arg="ecli"> (ECLI:IT:CASS:2022:32765CIV)</span><span style="display:none" data-role="content" data-arg="anno">2022</span><span class="label">,</span></span> </span> 
    <span class="chkcontent"><span class="label">udienza del</span>&nbsp;<span data-role="content" data-arg="datdec"><span style="font-weight:normal">19/10/2022</span></span><span class="label">,</span></span> 
    <span class="chkcontent"><span class="label">Presidente </span><span data-role="content" data-arg="presidente">PAOLITTO LIBERATO</span>&nbsp;</span> 
    <span class="chkcontent"><span class="label">Relatore </span><span data-role="content" data-arg="relatore">DELL'ORFANO ANTONELLA</span>&nbsp;</span> 
</a>

问题根源

  1. 页面内容动态加载:requests.get仅获取服务器返回的初始HTML,而包含PDF链接的<span class="toDocument pdf">标签是通过JavaScript动态渲染生成的,初始响应中不存在这些内容(从运行结果里的text2pdf标签内的空span可以验证)。
  2. XPath路径无效:你写的XPath路径对应渲染后的页面结构,但初始HTML中没有该结构,因此返回空列表。

解决方法

使用Selenium模拟浏览器加载页面,等待JavaScript完成渲染后再提取内容。步骤如下:

1. 安装依赖

pip install selenium

同时下载对应浏览器的驱动(如ChromeDriver),确保驱动版本与浏览器版本匹配,并将驱动路径配置到环境变量,或在代码中指定路径。

2. 修正后的代码

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from bs4 import BeautifulSoup

# 初始化浏览器驱动(这里以Chrome为例,需确保ChromeDriver路径正确)
driver = webdriver.Chrome()
driver.get('http://www.italgiure.giustizia.it/sncass/')

# 等待目标元素加载完成,最多等待10秒
try:
    WebDriverWait(driver, 10).until(
        EC.presence_of_element_located((By.CLASS_NAME, 'toDocument.pdf'))
    )
    # 获取渲染后的页面源码
    page_source = driver.page_source
    soup = BeautifulSoup(page_source, 'html.parser')
    
    # 提取所有PDF链接
    pdf_spans = soup.find_all('span', class_='toDocument pdf')
    pdf_links = []
    for span in pdf_spans:
        link = span.get('data-arg')
        # 拼接完整URL
        full_link = f'http://www.italgiure.giustizia.it{link}'
        pdf_links.append(full_link)
    
    print("获取到的PDF链接:")
    for link in pdf_links:
        print(link)
        
finally:
    # 关闭浏览器
    driver.quit()

代码说明

  • 用WebDriverWait等待目标元素出现,确保页面渲染完成。
  • 提取到data-arg中的相对路径后,拼接成完整的PDF URL。
  • 最后关闭浏览器驱动,避免资源占用。

内容的提问来源于stack exchange,提问作者Praveen Bushipaka

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.13 22:45:37