如何通过Scopus API提取指定年份前作者的总被引次数?
问题
我想通过Scopus提取某作者的被引次数,网页端能查看作者的被引概览,还能按特定年份筛选(比如<2016、<2019等)的被引数据。请问能不能通过API提取这类信息?
我写了一段Python代码,但只拿到了该作者2013年之前发表文章的Scopus ID,我需要的是这些文章的总被引次数总和。代码如下:
resp = requests.get( "http://api.elsevier.com/content/search/scopus", params={ "query": "au-id(34167797100) AND PUBYEAR < 2013", "start": "0", "count": "10", }, headers={"X-ELS-APIKey": key, "Accept": "application/json"}, ) #alle papers zum author holen data = resp.json() total_results = data["search-results"]["opensearch:totalResults"] results = [] for i in range(0, int(total_results), 25): resp = requests.get( "http://api.elsevier.com/content/search/scopus", params={ "query": "au-id(34167797100) AND PUBYEAR < 2013", "start": i, "count": "25", }, headers={"X-ELS-APIKey": key, "Accept": "application/json"}, ) results.extend(resp.json()["search-results"]["entry"]) liste_scopus_id = [] Path("resp.json").write_text(dumps(results, indent=4)) for result in results: liste_scopus_id.append(result["dc:identifier"].split(":")[1])
解决方案
Scopus API完全支持提取这类分年份的被引统计,有两种高效方法可以实现你的需求:
方法一:直接调用作者档案API获取分年被引数据
这种方法只需要一次请求就能拿到作者所有年份的被引统计,不用遍历每篇文章,效率最高。
import requests from json import dumps key = "你的API密钥" author_id = "34167797100" target_year = 2013 # 要筛选的年份阈值 # 请求作者档案,指定获取指标数据 resp = requests.get( f"https://api.elsevier.com/content/author/author_id/{author_id}", headers={"X-ELS-APIKey": key, "Accept": "application/json"}, params={"view": "metrics"} ) data = resp.json() # 提取分年份的被引记录 citation_year_data = data["author-retrieval-response"][0]["coredata"]["citation-counts"]["citation-count"] # 累加目标年份之前的被引次数 total_citations = 0 for item in citation_year_data: year = int(item["@year"]) if year < target_year: total_citations += int(item["$"]) print(f"该作者2013年之前发表文章的总被引次数:{total_citations}")
方法二:优化原搜索接口,直接返回被引次数
如果你想基于原有的搜索逻辑修改,可以在请求时指定返回citedby-count字段,这样每篇文章的被引次数会直接出现在搜索结果里,不用单独获取ID再查询。
import requests from pathlib import Path from json import dumps key = "你的API密钥" author_id = "34167797100" target_year = 2013 # 首次请求获取总结果数 resp = requests.get( "http://api.elsevier.com/content/search/scopus", params={ "query": f"au-id({author_id}) AND PUBYEAR < {target_year}", "start": "0", "count": "10", "field": "citedby-count" # 指定返回被引数字段 }, headers={"X-ELS-APIKey": key, "Accept": "application/json"}, ) data = resp.json() total_results = int(data["search-results"]["opensearch:totalResults"]) total_citations = 0 # 分页请求并累加被引次数 for i in range(0, total_results, 25): resp = requests.get( "http://api.elsevier.com/content/search/scopus", params={ "query": f"au-id({author_id}) AND PUBYEAR < {target_year}", "start": str(i), "count": "25", "field": "citedby-count" }, headers={"X-ELS-APIKey": key, "Accept": "application/json"}, ) entries = resp.json()["search-results"]["entry"] for entry in entries: # 处理无被引数据的文章,默认加0 total_citations += int(entry.get("citedby-count", 0)) print(f"该作者2013年之前发表文章的总被引次数:{total_citations}")
注意事项
- 确认你的API密钥有对应权限,部分指标接口需要机构授权才能访问
- 分页请求时遵守Scopus API的调用频率限制,避免触发限流
- 优先选择方法一,减少请求次数,提升效率
内容的提问来源于stack exchange,提问作者Hansi Hubertus Beidlus
相关产品推荐
相关产品推荐

