You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何从指定网页区域精准下载欧盟委员会提案PDF文件

问题描述

需从指定欧盟议会页面的特定区域仅下载欧盟委员会提案对应的PDF文件,现有代码会抓取所有.EN.pdf后缀的文件,包含不需要的后续跟进文档,需优化实现精准抓取或文件区分。

解决方案

思路1:基于HTML特征精准筛选

从目标PDF的HTML特征来看,目标链接要么带有class="externalDocument"属性,要么链接文本符合COM(XXXX)XXXX格式(如COM(2020)0791)。结合这两个特征可以精准过滤目标文件,同时也可以通过定位页面中"委员会提案"的专属区块进一步缩小范围。

思路2:文件名标记区分

如果无法精准定位区域,可在下载时给目标文件添加专属前缀(如TARGET_),方便后续快速区分目标提案和其他跟进文档。

修改后的代码

import os
import re
import requests
from urllib.parse import urljoin
from bs4 import BeautifulSoup

urls = ["http://www.europarl.europa.eu/oeil/FindByProcnum.do?lang=en&procnum=OLP/2020/0350", 
        "http://www.europarl.europa.eu/oeil/FindByProcnum.do?lang=en&procnum=OLP/2012/0299", 
        "http://www.europarl.europa.eu/oeil/FindByProcnum.do?lang=en&procnum=OLP/2013/0092"]

folder_location = r'C:\Users\myname\Documents\R\webscraping'
if not os.path.exists(folder_location):
    os.mkdir(folder_location)

# 匹配欧盟委员会提案编号的正则
com_proposal_pattern = re.compile(r'COM\(\d{4}\)\d{4}', re.IGNORECASE)

for url in urls:
    response = requests.get(url)
    soup = BeautifulSoup(response.text, "html.parser")
    
    # 方法1:筛选带externalDocument类且文本符合提案编号格式的PDF链接
    for link in soup.select("a.externalDocument[href$='EN.pdf']"):
        clean_text = link.get_text(strip=True)
        if com_proposal_pattern.match(clean_text):
            # 给目标文件名添加TARGET_前缀,方便区分
            filename = os.path.join(folder_location, f"TARGET_{link['href'].split('/')[-1]}")
            with open(filename, 'wb') as f:
                f.write(requests.get(urljoin(url, link['href'])).content)
            print(f"已下载目标提案:{filename}")
    
    # 方法2:定位页面中"Commission Proposal"区块内的PDF(可选,更精准)
    for proposal_header in soup.find_all('h3', string=lambda text: text and 'Commission Proposal' in text):
        # 获取标题所在的父容器,仅抓取该容器内的PDF
        container = proposal_header.find_parent()
        if container:
            for link in container.select("a[href$='EN.pdf']"):
                filename = os.path.join(folder_location, f"TARGET_{link['href'].split('/')[-1]}")
                with open(filename, 'wb') as f:
                    f.write(requests.get(urljoin(url, link['href'])).content)
            print(f"已下载指定区域内的提案文件")

代码说明

  • 加入正则表达式验证链接文本,确保仅抓取符合欧盟委员会提案编号格式的文件
  • 通过a.externalDocument选择器筛选目标链接,排除其他无关PDF
  • 下载时给文件名添加TARGET_前缀,直观区分目标文件和后续跟进文档
  • 新增区块定位逻辑,可精准抓取页面中"Commission Proposal"板块内的所有PDF,进一步避免误抓

内容的提问来源于stack exchange,提问作者Cesare

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.06 16:20:46