You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python爬取IEEE Xplore文献失败,请求技术协助

解决IEEE Xplore网页爬取无结果及错误页面问题

问题说明

需通过网页爬取实现IEEE Xplore技术情报自动化追踪,目标是获取2017年之后发表、包含“deep learning”和“radar”关键词的文献的年份、标题及摘要。现有代码执行后返回错误页面,articles列表为空。

用户原代码

import requests
from bs4 import BeautifulSoup

# 定义包含"Deep Learning"和"Radar"关键词的搜索URL
url = "https://ieeexplore.ieee.org/search/searchresult.jsp?action=search&newsearch=true&matchBoolean=true&queryText=(%22All%20Metadata%22:Deep%20learning)%20AND%20(%22All%20Metadata%22:radar)&ranges=2017_2023_Year"

# 发送GET请求并获取页面HTML内容
response = requests.get(url)
html = response.content

# 用BeautifulSoup解析HTML以提取所需信息
soup = BeautifulSoup(html, "html.parser")
articles = soup.select(".List-results-items")

# 遍历所有搜索结果,提取发表年份、标题和摘要
for article in articles:
    year = article.select_one(".publication-year").text.strip()
    title = article.select_one(".title").text.strip()
    abstract = article.select_one(".description").text.strip()

    # 打印结果
    print(f"发表年份: {year}")
    print(f"标题: {title}")
    print(f"摘要: {abstract}")
    print("\n")

遇到的错误

请求返回错误页面,解析后的内容如下:

<html><head><title>Error</title></head><body>
<html><head><title>Error</title></head><body>
An error occurred while processing your request.<p>
Reference #30.2e747e68.1677864420.73bb08c5
</p></body></html>
</p></body></html>

解决方案

1. 绕过反爬机制:添加请求头

IEEE Xplore会拦截无合法请求头的爬虫请求,需添加User-Agent模拟浏览器访问,同时可加入Accept等头信息提升请求合法性。

修正后的基础爬取代码:

import requests
from bs4 import BeautifulSoup

headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36",
    "Accept": "text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,*/*;q=0.8"
}

url = "https://ieeexplore.ieee.org/search/searchresult.jsp?action=search&newsearch=true&matchBoolean=true&queryText=(%22All%20Metadata%22:Deep%20learning)%20AND%20(%22All%20Metadata%22:radar)&ranges=2017_2023_Year"

# 带请求头发送请求
response = requests.get(url, headers=headers)
response.raise_for_status()  # 检查请求是否成功

soup = BeautifulSoup(response.text, "html.parser")
articles = soup.select(".List-results-items")

if not articles:
    print("未找到结果,可能需要处理分页或进一步绕过反爬")
else:
    for article in articles:
        year = article.select_one(".publication-year").text.strip() if article.select_one(".publication-year") else "未知"
        title = article.select_one(".title a").text.strip() if article.select_one(".title a") else "无标题"
        # 列表页不直接显示摘要,需进入详情页抓取
        abstract_link = article.select_one(".title a")["href"]
        abstract_url = f"https://ieeexplore.ieee.org{abstract_link}"
        
        # 抓取详情页的摘要
        abstract_response = requests.get(abstract_url, headers=headers)
        abstract_soup = BeautifulSoup(abstract_response.text, "html.parser")
        abstract = abstract_soup.select_one(".abstract-text").text.strip() if abstract_soup.select_one(".abstract-text") else "无摘要"
        
        print(f"发表年份: {year}")
        print(f"标题: {title}")
        print(f"摘要: {abstract}")
        print("\n")

2. 使用官方API(推荐方案)

直接爬取网页易触发反爬,且维护成本高。IEEE提供官方API,可稳定获取文献数据,需先申请API密钥。

核心逻辑示例:

import requests

API_KEY = "你的API密钥"
url = "https://ieeexploreapi.ieee.org/api/v1/search/articles"

params = {
    "apikey": API_KEY,
    "querytext": "\"Deep learning\" AND \"radar\"",
    "publication_year": "2017-2024",
    "max_records": 100,
    "start_record": 1,
    "fields": "article_title,publication_year,abstract"
}

response = requests.get(url, params=params)
data = response.json()

for article in data["articles"]:
    print(f"发表年份: {article.get('publication_year', '未知')}")
    print(f"标题: {article.get('article_title', '无标题')}")
    print(f"摘要: {article.get('abstract', '无摘要')}")
    print("\n")

说明:API请求需遵循调用频率限制,可通过start_record参数实现分页获取所有结果。

内容的提问来源于stack exchange,提问作者Pierrick Richard

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.29 23:20:32