You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

网站改版后,如何从新加坡国家档案馆演讲搜索页爬取指定演讲者PDF?

爬取Amy Khor演讲PDF文件的解决方案(适配网站改版)

问题描述

我需要从页面爬取演讲者Amy Khor的PDF文件,但该网站已改版,之前基于requests和BeautifulSoup的Python代码无法正常运行,请问该如何实现需求?
目标页面地址:https://www.nas.gov.sg/archivesonline/speeches/search-result?search-type=advanced&speaker=Amy+Khor

旧代码(已失效)

import requests
from bs4 import BeautifulSoup

url = 'http://www.nas.gov.sg/archivesonline/speeches/search-result?search-type=advanced&speaker='

search_term = 'Amy+Khor'

data = {
    'keywords': search_term,
    'search-type': 'basic',
    'keywords-type': 'all',
    'page-num': 1
}

soup = BeautifulSoup(requests.post(url, data=data).text, 'lxml')

cnt = 1
while True:

    print()
    print('Page no. {}'.format(cnt))
    print('-' * 80)

    for a in soup.select('a[href$=".pdf"]'):
        print(a['href'])

    if soup.select_one('span.next-10'):
        data['page-num'] += 10
        cnt += 1
        soup = BeautifulSoup(requests.post(url, data=data).text, 'lxml')
    else:
        break

解决方案

网站改版后采用了动态数据加载(AJAX接口),不再通过表单POST返回HTML页面,需要直接调用后端API获取数据:

1. 分析真实数据接口

通过浏览器开发者工具(F12)的Network标签抓包,发现列表数据来自https://www.nas.gov.sg/api/v2/archivesonline/search接口,请求方式为POST,需携带JSON格式的请求参数。

2. 适配后的代码实现

import requests
import json

# 目标API接口地址
api_url = "https://www.nas.gov.sg/api/v2/archivesonline/search"
# 请求头,模拟浏览器访问避免被拦截
headers = {
    "Content-Type": "application/json",
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36"
}

# 初始请求参数:指定演讲者、分页起始位置、每页数量
payload = {
    "query": {
        "bool": {
            "must": [
                {"term": {"speaker.keyword": "Amy Khor"}}
            ]
        }
    },
    "from": 0,
    "size": 10,
    "sort": [{"speech_date": {"order": "desc"}}]
}

all_pdf_links = []
current_page = 1

while True:
    print(f"正在爬取第 {current_page} 页...")
    # 发送POST请求获取数据
    response = requests.post(api_url, headers=headers, data=json.dumps(payload))
    response.raise_for_status()  # 捕获请求错误
    result = response.json()
    
    # 提取当前页的PDF链接
    hits = result.get("hits", {}).get("hits", [])
    if not hits:
        break  # 无更多数据,终止循环
    
    for item in hits:
        pdf_url = item.get("_source", {}).get("pdf_url")
        if pdf_url:
            all_pdf_links.append(pdf_url)
            print(pdf_url)
    
    # 更新分页参数,获取下一页数据
    payload["from"] += payload["size"]
    current_page += 1

print(f"\n爬取完成,共找到 {len(all_pdf_links)} 个PDF文件")

# 可选:自动下载所有PDF文件
# import os
# save_folder = "Amy_Khor_Speeches"
# os.makedirs(save_folder, exist_ok=True)
# for index, url in enumerate(all_pdf_links, 1):
#     print(f"正在下载第 {index} 个文件...")
#     pdf_response = requests.get(url, headers=headers)
#     with open(os.path.join(save_folder, f"speech_{index}.pdf"), "wb") as f:
#         f.write(pdf_response.content)

3. 注意事项

  • 必须携带User-Agent请求头,否则可能被网站拦截
  • 控制请求频率,可添加time.sleep(1)避免给服务器造成过大压力
  • 遵守网站的robots.txt规则,合理爬取数据

内容的提问来源于stack exchange,提问作者user3081750

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.25 17:18:26