如何爬取含on-click按钮网站的PDF链接?BeautifulSoup/lxml遇阻求助
问题描述
尝试用BeautifulSoup和lxml爬取http://www.italgiure.giustizia.it/sncass/的PDF文件,无法获取<span class="toDocument pdf">标签data-arg属性中的PDF链接,代码返回空列表。相关代码、运行结果及渲染后的HTML片段如下:
原代码
import requests from bs4 import BeautifulSoup from lxml import html r= requests.get('http://www.italgiure.giustizia.it/sncass/') soup = BeautifulSoup(r.text, 'html.parser') pdf_list = soup.find_all('a') print(pdf_list) search_html = html.fromstring(r.text) page_link = search_html.xpath('//*[@id="contentData"]/div[2]/div[1]/div/h3/a/span[1]/span') print(page_link)
运行结果
[<a href="accessibilita.html" style="text-decoration:none;font-size:80%;color:white" tabindex="0">Accessibilità</a>, <a accesskey="r" name="results" onclick="$(this).next().focus();" tabindex="-2" title="contenuto"></a>, <a accesskey="1" name="card" onclick="$(this).next().focus();" tabindex="-2" title="documento"></a>, <a class="text2pdf" href="javascript:void(0)" onclick="toTargetDoc($('.toDocument.pdf',$(this)).attr('data-arg'), this)" style="text-decoration:none;color:#440;" tabindex="0"> <span data-arg="filename" data-role="content" title="pdf"></span> <span class="chkcontent"><span class="label">Sez.</span> <span class="risultato" data-arg="szdec" data-role="content"></span> <span class="risultato" data-arg="kind" data-role="content"></span><span class="chkcontent"> - <span class="risultato" data-arg="ssz" data-role="content"></span></span><span class="label">,</span> </span> <span data-arg="tipoprov" data-role="content"></span> <span class="chkcontent"><span class="chkcontent"><span class="label">n.</span><span data-arg="numcard" data-role="content"></span></span><span data-arg="numdec" data-role="content" style="display:none"></span><span data-arg="numdep" data-role="content" style="display:none"></span> <span class="chkcontent"><span class="label"> del </span><span data-arg="datdep" data-role="content"></span><span data-arg="ecli" data-role="content" style="font-weight:normal"></span><span data-arg="anno" data-role="content" style="display:none"></span><span class="label">,</span></span> </span> <span class="chkcontent"><span class="label">udienza del</span> <span data-arg="datdec" data-role="content"></span><span class="label">,</span></span> <span class="chkcontent"><span class="label">Presidente </span><span data-arg="presidente" data-role="content"></span> </span> <span class="chkcontent"><span class="label">Relatore </span><span data-arg="relatore" data-role="content"></span> </span> </a>, <a class="text2ocr" href="javascript:void(0)" onclick="toTargetText($('.toDocument.txt',$(this)).attr('data-arg'))" style="text-decoration:none;color:#440;" tabindex="0"> <span data-arg="testoocr" data-role="content" title="testo ocr"></span> <span data-arg="ocr" data-role="datasubset"> <span data-arg="ocr" data-role="multivaluedcontent">snippet</span> </span> </a>, <a href="http://www.italgiure.giustizia.it" style="color:white;" tabindex="0">ItalgiureWeb</a>] []
渲染后的目标HTML片段
<a href="javascript:void(0)" tabindex="0" onclick="toTargetDoc($('.toDocument.pdf',$(this)).attr('data-arg'), this)" style="text-decoration:none;color:#440;" class="text2pdf"> <span data-role="content" data-arg="filename" title="pdf"> <span class="toDocument pdf" data-arg="/xway/application/nif/clean/hc.dll%3Fverbo%3Dattach%26db%3Dsnciv%26id%3D./20221107/snciv@s50@a2022@n32765@tO.clean.pdf"> <img class="rowIcon" alt="formato pdf" src="pix/pdf.png"> </span> </span> <span class="chkcontent"><span class="label">Sez.</span> <span class="risultato" data-role="content" data-arg="szdec">QUINTA</span> <span class="risultato" data-role="content" data-arg="kind">CIVILE</span><span class="label">,</span> </span> <span data-role="content" data-arg="tipoprov">Ordinanza</span> <span class="chkcontent"><span class="chkcontent"><span class="label">n.</span><span data-role="content" data-arg="numcard">32765</span></span><span style="display:none" data-role="content" data-arg="numdec">32765</span><span style="display:none" data-role="content" data-arg="numdep"></span> <span class="chkcontent"><span class="label"> del </span><span data-role="content" data-arg="datdep">07/11/2022</span><span style="font-weight:normal" data-role="content" data-arg="ecli"> (ECLI:IT:CASS:2022:32765CIV)</span><span style="display:none" data-role="content" data-arg="anno">2022</span><span class="label">,</span></span> </span> <span class="chkcontent"><span class="label">udienza del</span> <span data-role="content" data-arg="datdec"><span style="font-weight:normal">19/10/2022</span></span><span class="label">,</span></span> <span class="chkcontent"><span class="label">Presidente </span><span data-role="content" data-arg="presidente">PAOLITTO LIBERATO</span> </span> <span class="chkcontent"><span class="label">Relatore </span><span data-role="content" data-arg="relatore">DELL'ORFANO ANTONELLA</span> </span> </a>
问题根源
- 页面内容动态加载:
requests.get仅获取服务器返回的初始HTML,而包含PDF链接的<span class="toDocument pdf">标签是通过JavaScript动态渲染生成的,初始响应中不存在这些内容(从运行结果里的text2pdf标签内的空span可以验证)。 - XPath路径无效:你写的XPath路径对应渲染后的页面结构,但初始HTML中没有该结构,因此返回空列表。
解决方法
使用Selenium模拟浏览器加载页面,等待JavaScript完成渲染后再提取内容。步骤如下:
1. 安装依赖
pip install selenium
同时下载对应浏览器的驱动(如ChromeDriver),确保驱动版本与浏览器版本匹配,并将驱动路径配置到环境变量,或在代码中指定路径。
2. 修正后的代码
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from bs4 import BeautifulSoup # 初始化浏览器驱动(这里以Chrome为例,需确保ChromeDriver路径正确) driver = webdriver.Chrome() driver.get('http://www.italgiure.giustizia.it/sncass/') # 等待目标元素加载完成,最多等待10秒 try: WebDriverWait(driver, 10).until( EC.presence_of_element_located((By.CLASS_NAME, 'toDocument.pdf')) ) # 获取渲染后的页面源码 page_source = driver.page_source soup = BeautifulSoup(page_source, 'html.parser') # 提取所有PDF链接 pdf_spans = soup.find_all('span', class_='toDocument pdf') pdf_links = [] for span in pdf_spans: link = span.get('data-arg') # 拼接完整URL full_link = f'http://www.italgiure.giustizia.it{link}' pdf_links.append(full_link) print("获取到的PDF链接:") for link in pdf_links: print(link) finally: # 关闭浏览器 driver.quit()
代码说明
- 用
WebDriverWait等待目标元素出现,确保页面渲染完成。 - 提取到
data-arg中的相对路径后,拼接成完整的PDF URL。 - 最后关闭浏览器驱动,避免资源占用。
内容的提问来源于stack exchange,提问作者Praveen Bushipaka
相关产品推荐
相关产品推荐

