You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python可靠下载1969年意大利《Gazzetta Ufficiale》PDF?

程序化下载1969年意大利《Gazzetta Ufficiale – Serie Generale》PDF文件的问题

需求背景

为撰写学术文章,需批量下载1969年意大利《Gazzetta Ufficiale – Serie Generale》的"pubblicazione completa non certificata"格式PDF文件。目标网站提供1946-1985年的PDF存档索引:

  • 年份索引(1946–1985):https://www.gazzettaufficiale.it/ricerca/pdf/foglio_ordinario2/2/0/0?reset=true
  • 1969年Serie Generale存档列表:https://www.gazzettaufficiale.it/ricercaArchivioCompleto/serie_generale/1969
  • 详情页示例:.../gazzetta/serie_generale/caricaDettaglio?dataPubblicazioneGazzetta=1969-01-02&numeroGazzetta=1

已尝试的失败方法

  1. Selenium:导航至年份选择器点击1969,但频繁超时;年份链接/输入框存在缺失、隐藏于iframe或被横幅遮挡等问题,切换框架、注入JS的方案仍不稳定。
  2. Requests + BeautifulSoup:年份网格页曾出现<a class="download_pdf" href="/do/gazzetta/downloadPdf?...">Download PDF</a>这类直接下载链接,但实时会话中无该锚点,抓取结果为0个链接。
  3. 手动构建下载URL:通过存档列表的日期和期号构建下载链接/do/gazzetta/downloadPdf?dataPubblicazioneGazzetta=19690102&numeroGazzetta=1&tipoSerie=SG&tipoSupplemento=GU&numeroSupplemento=0&progressivo=0&edizione=0&estensione=pdf,返回HTML而非PDF,内容提示“Il pdf selezionato non è stato trovato”,保存为.pdf文件会被Acrobat判定损坏。
  4. 请求详情页:获取详情页后,查找a.download_pdf或含“Scarica il PDF”/“pubblicazione completa non certificata”的锚点,1969年的页面中始终无此类链接,无法获取有效downloadPdf URL。

可复现代码

import requests
from bs4 import BeautifulSoup
from urllib.parse import urljoin, urlparse, parse_qs

BASE = "https://www.gazzettaufficiale.it"
YEAR = 1969
YEAR_URL = f"{BASE}/ricercaArchivioCompleto/serie_generale/{YEAR}"

s = requests.Session()
s.headers.update({
    "User-Agent": "Mozilla/5.0",
    "Accept": "text/html,application/xhtml+xml,application/xml;q=0.9,*/*;q=0.8",
    "Referer": BASE,
})

# 1) Collect detail pages (date + issue number)
r = s.get(YEAR_URL, timeout=60)
r.raise_for_status()
soup = BeautifulSoup(r.text, "html.parser")
details = []
for a in soup.find_all("a", href=True):
    href = a["href"]
    if ("/gazzetta/serie_generale/caricaDettaglio" in href
        and "dataPubblicazioneGazzetta=" in href
        and "numeroGazzetta=" in href):
        details.append(urljoin(BASE, href))

print("Detail pages found:", len(details))
print("Sample:", details[:3])

# 2) For one detail page, try to discover a real "download PDF" link
detail_url = details[0]
r = s.get(detail_url, timeout=60, headers={"Referer": YEAR_URL})
r.raise_for_status()
soup = BeautifulSoup(r.text, "html.parser")

# Try common selectors / texts
dl = (soup.select_one('a.download_pdf[href]')
      or soup.select_one('a[href*="/do/gazzetta/downloadPdf"]'))
if not dl:
    for a in soup.find_all("a", href=True):
        if "scarica il pdf" in (a.get_text() or "").lower():
            dl = a
            break

print("Download link found on detail page?", bool(dl))
if dl:
    print("Download href:", urljoin(BASE, dl["href"]))

代码运行输出

Detail pages found: 264
Sample: [https://www.gazzettaufficiale.it/gazzetta/serie_generale/caricaDettaglio?dataPubblicazioneGazzetta=1969-01-02&numeroGazzetta=1, ...]
Download link found on detail page? False

寻求帮助

求可行的解决方案或可用的最小示例。

内容的提问来源于stack exchange,提问作者Mark

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.12 14:14:58