如何用Python从Google Scholar搜索结果爬取完整论文引用格式
Google Scholar 完整APA引用爬取方案
核心逻辑
你需要的APA引用内容未内置在搜索列表页的静态HTML中,需要调用谷歌学术公开的引用生成接口单独请求获取,具体实现步骤如下:
- 从搜索结果页提取每篇论文的唯一ID:每个论文条目对应的
div.gs_r节点带有data-cid属性,属性值为谷歌学术分配的论文唯一标识,可直接提取。 - 构造引用接口请求URL:标准接口格式为
https://scholar.google.com/scholar?q=info:{论文ID}:scholar.google.com/&output=cite&hl=fr,hl参数和你现有请求的语言参数保持一致即可,将该URL拼接ScraperAPI前缀后发起请求。 - 解析返回内容提取APA引用:接口返回的HTML中,第一个
div.gs_citr节点的文本内容就是标准APA格式引用,直接提取即可。
注意事项
- 每获取1篇论文的引用会消耗1次ScraperAPI请求额度,你可按需添加筛选条件(例如仅爬取被引量大于指定值的论文),节省免费额度。
- 可适当添加请求间隔,避免触发谷歌学术的频次限制。
修改后的代码示例
import requests import numpy as np import pandas as pd import re from bs4 import BeautifulSoup import time APIKEY = "????????????????????" BASE_URL = f"http://api.scraperapi.com?api_key={APIKEY}&url=" def scraper_api(query, n_pages, get_apa=True, min_cite_for_apa=0): """Uses scraperAPI to scrape Google Scholar for papers' Title, Year, Citations, Cited By url and APA citation returns a dataframe --------------------------- parameters: query: in the following format "automation+container+terminal" n_pages: number of pages to scrape get_apa: whether to fetch APA citation, default True min_cite_for_apa: only fetch APA for papers with citations >= this value, default 0 --------------------------- returns: dataframe with paper info columns ---------------------------""" pages = np.arange(0,(n_pages*10),10) papers = [] for page in pages: print(f"Scraping page {int(page/10) + 1}") webpage = f"https://scholar.google.com/scholar?start={page}&q={query}&hl=fr&as_sdt=0,5" url = BASE_URL + webpage response = requests.get(url) soup = BeautifulSoup(response.content, "html.parser") # 遍历每篇论文的外层节点获取data-cid for paper_block in soup.find_all("div", class_="gs_r"): paper_id = paper_block.get("data-cid", "") paper = paper_block.find("div", class_="gs_ri") if not paper: continue # 原有逻辑:提取标题、年份、被引量、被引链接 # get the title of each paper title = paper.find("h3", class_="gs_rt").find("a").text if paper.find("h3", class_="gs_rt").find("a") else paper.find("h3", class_="gs_rt").find("span").text # get the year of publication of each paper txt_year = paper.find("div", class_="gs_a").text year = re.findall('[0-9]{4}', txt_year) year = list(map(int,year))[0] if year else 0 # get number of citations for each paper txt_cite = paper.find("div", class_="gs_fl").find_all("a")[2].string citations = 0 if txt_cite: cite_num = re.findall('[0-9]+', txt_cite) citations = list(map(int,cite_num))[0] if cite_num else 0 # get the "cited_by" url new_url = "no citations" urls = paper.find("div", class_="gs_fl").find_all(href=True) if urls: for url_item in urls: if "cites" in url_item["href"]: cited_url = url_item["href"] index1 = cited_url.index("?") url_slices = [] url_slices.append(cited_url[:index1+1]) url_slices.append(cited_url[index1+1:]) index_and = url_slices[1].index("&") url_slices.append(url_slices[1][:index_and+1]) url_slices.append(url_slices[1][index_and+1:]) del url_slices[1] new_url = "https://scholar.google.com.tw"+url_slices[0]+"start=00&hl=en&"+url_slices[2]+url_slices[1]+"scipsc=" break # 新增:提取APA引用 apa_citation = "" if get_apa and paper_id and citations >= min_cite_for_apa: cite_url = f"https://scholar.google.com/scholar?q=info:{paper_id}:scholar.google.com/&output=cite&hl=fr" full_cite_url = BASE_URL + cite_url cite_resp = requests.get(full_cite_url) cite_soup = BeautifulSoup(cite_resp.content, "html.parser") apa_node = cite_soup.find("div", class_="gs_citr") apa_citation = apa_node.text.strip() if apa_node else "获取失败" time.sleep(0.5) # 可自行调整间隔 # 存入结果 papers.append({ 'title': title, 'year': year, 'citations': citations, 'cited_by_url': new_url, 'apa_citation': apa_citation }) papers_df = pd.DataFrame(papers) return papers_df
内容的提问来源于stack exchange,提问作者jhupiterz
相关产品推荐
相关产品推荐

