如何用Python可靠下载1969年意大利《Gazzetta Ufficiale》PDF?
程序化下载1969年意大利《Gazzetta Ufficiale – Serie Generale》PDF文件的问题
需求背景
为撰写学术文章,需批量下载1969年意大利《Gazzetta Ufficiale – Serie Generale》的"pubblicazione completa non certificata"格式PDF文件。目标网站提供1946-1985年的PDF存档索引:
- 年份索引(1946–1985):
https://www.gazzettaufficiale.it/ricerca/pdf/foglio_ordinario2/2/0/0?reset=true - 1969年Serie Generale存档列表:
https://www.gazzettaufficiale.it/ricercaArchivioCompleto/serie_generale/1969 - 详情页示例:
.../gazzetta/serie_generale/caricaDettaglio?dataPubblicazioneGazzetta=1969-01-02&numeroGazzetta=1
已尝试的失败方法
- Selenium:导航至年份选择器点击1969,但频繁超时;年份链接/输入框存在缺失、隐藏于iframe或被横幅遮挡等问题,切换框架、注入JS的方案仍不稳定。
- Requests + BeautifulSoup:年份网格页曾出现
<a class="download_pdf" href="/do/gazzetta/downloadPdf?...">Download PDF</a>这类直接下载链接,但实时会话中无该锚点,抓取结果为0个链接。 - 手动构建下载URL:通过存档列表的日期和期号构建下载链接
/do/gazzetta/downloadPdf?dataPubblicazioneGazzetta=19690102&numeroGazzetta=1&tipoSerie=SG&tipoSupplemento=GU&numeroSupplemento=0&progressivo=0&edizione=0&estensione=pdf,返回HTML而非PDF,内容提示“Il pdf selezionato non è stato trovato”,保存为.pdf文件会被Acrobat判定损坏。 - 请求详情页:获取详情页后,查找
a.download_pdf或含“Scarica il PDF”/“pubblicazione completa non certificata”的锚点,1969年的页面中始终无此类链接,无法获取有效downloadPdf URL。
可复现代码
import requests from bs4 import BeautifulSoup from urllib.parse import urljoin, urlparse, parse_qs BASE = "https://www.gazzettaufficiale.it" YEAR = 1969 YEAR_URL = f"{BASE}/ricercaArchivioCompleto/serie_generale/{YEAR}" s = requests.Session() s.headers.update({ "User-Agent": "Mozilla/5.0", "Accept": "text/html,application/xhtml+xml,application/xml;q=0.9,*/*;q=0.8", "Referer": BASE, }) # 1) Collect detail pages (date + issue number) r = s.get(YEAR_URL, timeout=60) r.raise_for_status() soup = BeautifulSoup(r.text, "html.parser") details = [] for a in soup.find_all("a", href=True): href = a["href"] if ("/gazzetta/serie_generale/caricaDettaglio" in href and "dataPubblicazioneGazzetta=" in href and "numeroGazzetta=" in href): details.append(urljoin(BASE, href)) print("Detail pages found:", len(details)) print("Sample:", details[:3]) # 2) For one detail page, try to discover a real "download PDF" link detail_url = details[0] r = s.get(detail_url, timeout=60, headers={"Referer": YEAR_URL}) r.raise_for_status() soup = BeautifulSoup(r.text, "html.parser") # Try common selectors / texts dl = (soup.select_one('a.download_pdf[href]') or soup.select_one('a[href*="/do/gazzetta/downloadPdf"]')) if not dl: for a in soup.find_all("a", href=True): if "scarica il pdf" in (a.get_text() or "").lower(): dl = a break print("Download link found on detail page?", bool(dl)) if dl: print("Download href:", urljoin(BASE, dl["href"]))
代码运行输出
Detail pages found: 264 Sample: [https://www.gazzettaufficiale.it/gazzetta/serie_generale/caricaDettaglio?dataPubblicazioneGazzetta=1969-01-02&numeroGazzetta=1, ...] Download link found on detail page? False
寻求帮助
求可行的解决方案或可用的最小示例。
内容的提问来源于stack exchange,提问作者Mark
相关产品推荐
相关产品推荐

