You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python从Google Scholar搜索结果爬取完整论文引用格式

Google Scholar 完整APA引用爬取方案

核心逻辑

你需要的APA引用内容未内置在搜索列表页的静态HTML中,需要调用谷歌学术公开的引用生成接口单独请求获取,具体实现步骤如下:

  • 从搜索结果页提取每篇论文的唯一ID:每个论文条目对应的div.gs_r节点带有data-cid属性,属性值为谷歌学术分配的论文唯一标识,可直接提取。
  • 构造引用接口请求URL:标准接口格式为 https://scholar.google.com/scholar?q=info:{论文ID}:scholar.google.com/&output=cite&hl=fr,hl参数和你现有请求的语言参数保持一致即可,将该URL拼接ScraperAPI前缀后发起请求。
  • 解析返回内容提取APA引用:接口返回的HTML中,第一个div.gs_citr节点的文本内容就是标准APA格式引用,直接提取即可。

注意事项

  • 每获取1篇论文的引用会消耗1次ScraperAPI请求额度,你可按需添加筛选条件(例如仅爬取被引量大于指定值的论文),节省免费额度。
  • 可适当添加请求间隔,避免触发谷歌学术的频次限制。

修改后的代码示例

import requests
import numpy as np
import pandas as pd
import re
from bs4 import BeautifulSoup
import time

APIKEY = "????????????????????"
BASE_URL = f"http://api.scraperapi.com?api_key={APIKEY}&url="

def scraper_api(query, n_pages, get_apa=True, min_cite_for_apa=0):
    """Uses scraperAPI to scrape Google Scholar for 
    papers' Title, Year, Citations, Cited By url and APA citation returns a dataframe
    ---------------------------
    parameters:
    query: in the following format "automation+container+terminal"
    n_pages: number of pages to scrape
    get_apa: whether to fetch APA citation, default True
    min_cite_for_apa: only fetch APA for papers with citations >= this value, default 0
    ---------------------------
    returns:
    dataframe with paper info columns
    ---------------------------"""

    pages = np.arange(0,(n_pages*10),10)
    papers = []
    for page in pages:
        print(f"Scraping page {int(page/10) + 1}")
        webpage = f"https://scholar.google.com/scholar?start={page}&q={query}&hl=fr&as_sdt=0,5"
        url = BASE_URL + webpage
        response = requests.get(url)
        soup = BeautifulSoup(response.content, "html.parser")
        
        # 遍历每篇论文的外层节点获取data-cid
        for paper_block in soup.find_all("div", class_="gs_r"):
            paper_id = paper_block.get("data-cid", "")
            paper = paper_block.find("div", class_="gs_ri")
            if not paper:
                continue
            # 原有逻辑:提取标题、年份、被引量、被引链接
            # get the title of each paper
            title = paper.find("h3", class_="gs_rt").find("a").text if paper.find("h3", class_="gs_rt").find("a") else paper.find("h3", class_="gs_rt").find("span").text
            # get the year of publication of each paper
            txt_year = paper.find("div", class_="gs_a").text
            year = re.findall('[0-9]{4}', txt_year)
            year = list(map(int,year))[0] if year else 0
            # get number of citations for each paper
            txt_cite = paper.find("div", class_="gs_fl").find_all("a")[2].string
            citations = 0
            if txt_cite:
                cite_num = re.findall('[0-9]+', txt_cite)
                citations = list(map(int,cite_num))[0] if cite_num else 0
            # get the "cited_by" url
            new_url = "no citations"
            urls = paper.find("div", class_="gs_fl").find_all(href=True)
            if urls:
                for url_item in urls:
                    if "cites" in url_item["href"]:
                        cited_url = url_item["href"]
                        index1 = cited_url.index("?")
                        url_slices = []
                        url_slices.append(cited_url[:index1+1])
                        url_slices.append(cited_url[index1+1:])
                        index_and = url_slices[1].index("&")
                        url_slices.append(url_slices[1][:index_and+1])
                        url_slices.append(url_slices[1][index_and+1:])
                        del url_slices[1]
                        new_url = "https://scholar.google.com.tw"+url_slices[0]+"start=00&hl=en&"+url_slices[2]+url_slices[1]+"scipsc="
                        break
            # 新增:提取APA引用
            apa_citation = ""
            if get_apa and paper_id and citations >= min_cite_for_apa:
                cite_url = f"https://scholar.google.com/scholar?q=info:{paper_id}:scholar.google.com/&output=cite&hl=fr"
                full_cite_url = BASE_URL + cite_url
                cite_resp = requests.get(full_cite_url)
                cite_soup = BeautifulSoup(cite_resp.content, "html.parser")
                apa_node = cite_soup.find("div", class_="gs_citr")
                apa_citation = apa_node.text.strip() if apa_node else "获取失败"
                time.sleep(0.5) # 可自行调整间隔
            # 存入结果
            papers.append({
                'title': title, 
                'year': year, 
                'citations': citations, 
                'cited_by_url': new_url,
                'apa_citation': apa_citation
            })
    papers_df = pd.DataFrame(papers)
    return papers_df

内容的提问来源于stack exchange,提问作者jhupiterz

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.02 00:45:03