如何去除Semantic Scholar合著者统计代码冗余输出并优化?
问题
现有一段基于Semantic Scholar API的Python代码,运行无报错但会输出大量冗余信息,希望仅保留TOP10合著者及其合著次数的指定格式输出(如下示例),同时需要简化优化代码,仅允许使用已导入的库和指定API链接。
期望输出格式:
('Diego Calvanese', 47) ('D. Lanti', 28) ('Martín Rezk', 21) ('Elem Güzel Kalayci', 18) ('B. Cogrel', 17) ('E. Botoeva', 16) ('E. Kharlamov', 16) ('I. Horrocks', 12) ('S. Brandt', 11) ('V. Ryzhikov', 11)
原代码:
import urllib.request import json from collections import Counter def count_coauthors(author_id): coauthors_dict = {} url_str = ('https://api.semanticscholar.org/graph/v1/author/47490276?fields=name,papers.authors') respons = urllib.request.urlopen(url_str) text = respons.read().decode() for line in respons: print(line.decode().rstip()) data = json.loads(text) print(type(data)) print(list(data.keys())) print(data["name"]) print(data["authorId"]) name = [] for lines in data["papers"]: for authors in lines["authors"]: name.append(authors.get("name")) print(name) count = dict() names = name for i in names: if i not in count: count[i] = 1 else: count[i] += 1 print(count) c = Counter(count) top = c.most_common(10) print(top) return coauthors_dict author_id = '47490276' cc = count_coauthors(author_id) top_coauthors = sorted(cc.items(), key=lambda item: item[1], reverse=True) for co_author in top_coauthors[:10]: print(co_author)
优化后的代码
import urllib.request import json from collections import Counter def count_coauthors(author_id): # 动态拼接API链接,替换硬编码的author_id url_str = f'https://api.semanticscholar.org/graph/v1/author/{author_id}?fields=name,papers.authors' response = urllib.request.urlopen(url_str) text = response.read().decode() data = json.loads(text) author_name = data["name"] coauthors = [] # 遍历所有论文的作者,排除自己 for paper in data["papers"]: for author in paper["authors"]: name = author.get("name") if name != author_name: coauthors.append(name) # 直接用Counter统计次数 coauthor_counts = Counter(coauthors) # 获取TOP10合著者 top_10 = coauthor_counts.most_common(10) # 按指定格式输出 for item in top_10: print(item) return coauthor_counts author_id = '47490276' count_coauthors(author_id)
优化说明
- 移除所有冗余
print语句,仅保留目标输出 - 修复API链接硬编码问题,改为动态传入
author_id拼接 - 增加排除作者自身的逻辑,避免无效统计
- 简化计数流程:直接用
Counter统计合著者列表,省去手动遍历计数的冗余代码 - 直接在函数内输出指定格式结果,无需返回后二次处理
- 修正变量名拼写错误(
respons改为response)
内容的提问来源于stack exchange,提问作者marlene
相关产品推荐
相关产品推荐

