You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

网页爬虫无法收集URL:司法网站PFD报告爬取代码适配求助

问题

我正在爬取某司法网站的预防未来死亡报告页面以创建数据库,原有Python代码因网站HTML结构完全变更而失效。新代码无法定位报告URL,尝试过将选择器替换为card__link和li均无效,当前代码如下:

import requests
from bs4 import BeautifulSoup
from tqdm import tqdm

base_url = 'https://www.judiciary.uk/prevention-of-future-death-reports/page/{}/'
page_count = 442

with requests.Session() as session:
    record_urls = [ ]
    for page in tqdm(range(1, page_count + 1)):
        url = base_url.format(page)
        try:
            response = session.get(url)
            response.raise_for_status()
            soup = BeautifulSoup(response.content, 'html.parser')
            record_urls += [h5.a['href'] for h5 in soup.find_all('h5', {'class': 'entry-title'})]
        except (requests.exceptions.RequestException, ValueError, AttributeError) as e:
            print(f"Failed to process page {page}: {e}")
    print(f"Collected {len(record_urls)} URLs")

需要找到当前包含报告URL的HTML元素,以收集约4420个URL。

解决方案

1. 定位URL元素的方法

打开目标分页页面,按F12调出开发者工具:

  • 找到任意报告标题的链接,右键选择「检查」
  • 可看到URL所在的a标签带有listing__link类,该标签直接包含报告标题,且属于ul.listing__items下的li.listing__item列表项

2. 修改后的代码

将原代码中定位元素的逻辑替换为寻找带有listing__link类的a标签,直接提取其href属性:

import requests
from bs4 import BeautifulSoup
from tqdm import tqdm

base_url = 'https://www.judiciary.uk/prevention-of-future-death-reports/page/{}/'
page_count = 442

with requests.Session() as session:
    record_urls = []
    for page in tqdm(range(1, page_count + 1)):
        url = base_url.format(page)
        try:
            response = session.get(url)
            response.raise_for_status()
            soup = BeautifulSoup(response.content, 'html.parser')
            # 定位所有带listing__link类的a标签,提取href
            record_urls += [link['href'] for link in soup.find_all('a', {'class': 'listing__link'})]
        except (requests.exceptions.RequestException, ValueError, AttributeError) as e:
            print(f"Failed to process page {page}: {e}")
    print(f"Collected {len(record_urls)} URLs")

说明

原尝试的card__link是旧结构或其他模块的类名,当前页面的报告列表使用listing__link类标识链接,修改后即可正确提取每页的10个报告URL,最终收集到约4420个链接。

内容的提问来源于stack exchange,提问作者Georgia Richards

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.22 01:47:10