Python爬取IEEE Xplore文献失败,请求技术协助
解决IEEE Xplore网页爬取无结果及错误页面问题
问题说明
需通过网页爬取实现IEEE Xplore技术情报自动化追踪,目标是获取2017年之后发表、包含“deep learning”和“radar”关键词的文献的年份、标题及摘要。现有代码执行后返回错误页面,articles列表为空。
用户原代码
import requests from bs4 import BeautifulSoup # 定义包含"Deep Learning"和"Radar"关键词的搜索URL url = "https://ieeexplore.ieee.org/search/searchresult.jsp?action=search&newsearch=true&matchBoolean=true&queryText=(%22All%20Metadata%22:Deep%20learning)%20AND%20(%22All%20Metadata%22:radar)&ranges=2017_2023_Year" # 发送GET请求并获取页面HTML内容 response = requests.get(url) html = response.content # 用BeautifulSoup解析HTML以提取所需信息 soup = BeautifulSoup(html, "html.parser") articles = soup.select(".List-results-items") # 遍历所有搜索结果,提取发表年份、标题和摘要 for article in articles: year = article.select_one(".publication-year").text.strip() title = article.select_one(".title").text.strip() abstract = article.select_one(".description").text.strip() # 打印结果 print(f"发表年份: {year}") print(f"标题: {title}") print(f"摘要: {abstract}") print("\n")
遇到的错误
请求返回错误页面,解析后的内容如下:
<html><head><title>Error</title></head><body> <html><head><title>Error</title></head><body> An error occurred while processing your request.<p> Reference #30.2e747e68.1677864420.73bb08c5 </p></body></html> </p></body></html>
解决方案
1. 绕过反爬机制:添加请求头
IEEE Xplore会拦截无合法请求头的爬虫请求,需添加User-Agent模拟浏览器访问,同时可加入Accept等头信息提升请求合法性。
修正后的基础爬取代码:
import requests from bs4 import BeautifulSoup headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36", "Accept": "text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,*/*;q=0.8" } url = "https://ieeexplore.ieee.org/search/searchresult.jsp?action=search&newsearch=true&matchBoolean=true&queryText=(%22All%20Metadata%22:Deep%20learning)%20AND%20(%22All%20Metadata%22:radar)&ranges=2017_2023_Year" # 带请求头发送请求 response = requests.get(url, headers=headers) response.raise_for_status() # 检查请求是否成功 soup = BeautifulSoup(response.text, "html.parser") articles = soup.select(".List-results-items") if not articles: print("未找到结果,可能需要处理分页或进一步绕过反爬") else: for article in articles: year = article.select_one(".publication-year").text.strip() if article.select_one(".publication-year") else "未知" title = article.select_one(".title a").text.strip() if article.select_one(".title a") else "无标题" # 列表页不直接显示摘要,需进入详情页抓取 abstract_link = article.select_one(".title a")["href"] abstract_url = f"https://ieeexplore.ieee.org{abstract_link}" # 抓取详情页的摘要 abstract_response = requests.get(abstract_url, headers=headers) abstract_soup = BeautifulSoup(abstract_response.text, "html.parser") abstract = abstract_soup.select_one(".abstract-text").text.strip() if abstract_soup.select_one(".abstract-text") else "无摘要" print(f"发表年份: {year}") print(f"标题: {title}") print(f"摘要: {abstract}") print("\n")
2. 使用官方API(推荐方案)
直接爬取网页易触发反爬,且维护成本高。IEEE提供官方API,可稳定获取文献数据,需先申请API密钥。
核心逻辑示例:
import requests API_KEY = "你的API密钥" url = "https://ieeexploreapi.ieee.org/api/v1/search/articles" params = { "apikey": API_KEY, "querytext": "\"Deep learning\" AND \"radar\"", "publication_year": "2017-2024", "max_records": 100, "start_record": 1, "fields": "article_title,publication_year,abstract" } response = requests.get(url, params=params) data = response.json() for article in data["articles"]: print(f"发表年份: {article.get('publication_year', '未知')}") print(f"标题: {article.get('article_title', '无标题')}") print(f"摘要: {article.get('abstract', '无摘要')}") print("\n")
说明:API请求需遵循调用频率限制,可通过start_record参数实现分页获取所有结果。
内容的提问来源于stack exchange,提问作者Pierrick Richard
相关产品推荐
相关产品推荐

