You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何通过EC Search API获取欧盟资助项目的目标信息?

解决方案:获取EC项目招标的目标信息

1. 优先检查Search API是否支持返回目标字段

先排查API返回的完整结构,当前代码只提取了keywords和content,可能目标字段藏在metadata或其他深层结构中:

  • 打印完整的result对象(比如print(json.dumps(result, indent=2))),查看是否有objectives、description或类似字段。
  • 尝试在API请求中添加字段指定参数,部分搜索API支持通过_source或fields参数强制返回特定字段,可在files中新增:
    "_source": ("blob", json.dumps(["metadata.objectives", "content"]), "application/json"),
    

2. 若API无法直接获取,通过网页抓取详情页内容

每个项目的详情页URL已包含在result["url"]中,可通过请求详情页并解析HTML提取目标板块,以下是修改后的代码示例:

首先安装依赖:

pip install beautifulsoup4

修改后的完整代码:

import json
import requests
import time
from bs4 import BeautifulSoup

api_url = "https://api.tech.ec.europa.eu/search-api/prod/rest/search"
params = {"apiKey": "SEDIA", "text": "***", "pageSize": "50", "pageNumber": "1"}

query = {
    "bool": {
        "must": [
            {"terms": {"type": ["1", "2"]}},
            {"terms": {"status": ["31094502"]}},
            {"term": {"programmePeriod": "2021 - 2027"}},
            {"terms": {"frameworkProgramme": ["43108390"]}},
        ]
    }
}

languages = ["en"]
sort = {"field": "sortStatus", "order": "ASC"}

page = 1
has_more_results = True

# 模拟浏览器请求头,避免被拦截
headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36"
}

while has_more_results:
    params["pageNumber"] = str(page)
    data = requests.post(
        api_url,
        params=params,
        files={
            "query": ("blob", json.dumps(query), "application/json"),
            "languages": ("blob", json.dumps(languages), "application/json"),
            "sort": ("blob", json.dumps(sort), "application/json"),
        },
    ).json()

    for result in data["results"]:
        keywords = result["metadata"]["keywords"][0] 
        content = result["content"]
        url = result["url"]
        
        print("Project Name:{}\nKeywords: {}".format(content, keywords))
        
        # 抓取详情页提取目标内容
        try:
            # 添加请求间隔,避免触发反爬
            time.sleep(1)
            detail_response = requests.get(url, headers=headers)
            detail_response.raise_for_status()
            soup = BeautifulSoup(detail_response.text, "html.parser")
            
            # 根据详情页实际HTML结构定位目标板块,需自行调整
            # 示例:匹配"Objectives"标题后的内容块
            objectives_title = soup.find("h3", string=lambda text: text and "Objectives" in text)
            if objectives_title:
                objectives_content = objectives_title.find_next_sibling(class_="ecl-u-mb-l").get_text(strip=True, separator="\n")
                print("Objectives:\n{}".format(objectives_content))
            else:
                print("Objectives section not found.")
        except Exception as e:
            print(f"Failed to fetch details: {str(e)}")
        
        print("-" * 50)

    has_more_results = len(data["results"]) == int(params["pageSize"])
    page += 1

注意:需通过浏览器开发者工具查看详情页的HTML结构,调整find方法的参数(比如标签类型、class名称)来精准定位目标板块。

3. 额外建议

  • 查看EC官方文档,确认是否存在专门获取项目详情的API端点(比如/projects/{project-id}),这类接口通常返回更完整的字段。
  • 若遇到反爬限制,可考虑使用代理IP或增加请求间隔,避免IP被封禁。

内容的提问来源于stack exchange,提问作者Furkan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.23 09:35:37