网页爬虫无法收集URL:司法网站PFD报告爬取代码适配求助
问题
我正在爬取某司法网站的预防未来死亡报告页面以创建数据库,原有Python代码因网站HTML结构完全变更而失效。新代码无法定位报告URL,尝试过将选择器替换为card__link和li均无效,当前代码如下:
import requests from bs4 import BeautifulSoup from tqdm import tqdm base_url = 'https://www.judiciary.uk/prevention-of-future-death-reports/page/{}/' page_count = 442 with requests.Session() as session: record_urls = [ ] for page in tqdm(range(1, page_count + 1)): url = base_url.format(page) try: response = session.get(url) response.raise_for_status() soup = BeautifulSoup(response.content, 'html.parser') record_urls += [h5.a['href'] for h5 in soup.find_all('h5', {'class': 'entry-title'})] except (requests.exceptions.RequestException, ValueError, AttributeError) as e: print(f"Failed to process page {page}: {e}") print(f"Collected {len(record_urls)} URLs")
需要找到当前包含报告URL的HTML元素,以收集约4420个URL。
解决方案
1. 定位URL元素的方法
打开目标分页页面,按F12调出开发者工具:
- 找到任意报告标题的链接,右键选择「检查」
- 可看到URL所在的
a标签带有listing__link类,该标签直接包含报告标题,且属于ul.listing__items下的li.listing__item列表项
2. 修改后的代码
将原代码中定位元素的逻辑替换为寻找带有listing__link类的a标签,直接提取其href属性:
import requests from bs4 import BeautifulSoup from tqdm import tqdm base_url = 'https://www.judiciary.uk/prevention-of-future-death-reports/page/{}/' page_count = 442 with requests.Session() as session: record_urls = [] for page in tqdm(range(1, page_count + 1)): url = base_url.format(page) try: response = session.get(url) response.raise_for_status() soup = BeautifulSoup(response.content, 'html.parser') # 定位所有带listing__link类的a标签,提取href record_urls += [link['href'] for link in soup.find_all('a', {'class': 'listing__link'})] except (requests.exceptions.RequestException, ValueError, AttributeError) as e: print(f"Failed to process page {page}: {e}") print(f"Collected {len(record_urls)} URLs")
说明
原尝试的card__link是旧结构或其他模块的类名,当前页面的报告列表使用listing__link类标识链接,修改后即可正确提取每页的10个报告URL,最终收集到约4420个链接。
内容的提问来源于stack exchange,提问作者Georgia Richards
相关产品推荐
相关产品推荐

