You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用SimpleHtmlDom爬取Google Scholar个人主页文献的数量限制问题

解决Google Scholar个人出版物全量爬取问题

我懂你现在的困扰——用SimpleHtmlDom只能抓到当前页面显示的内容,哪怕把pagesize调到100,超过这个数量的出版物还是拿不到。其实核心问题是Google Scholar的内容是分页加载的,咱们得通过循环遍历所有分页来获取全部内容。

核心思路

Google Scholar的分页是通过cstart参数控制的:

  • cstart=0:对应第1-100条(搭配pagesize=100)
  • cstart=100:对应第101-200条
  • cstart=200:对应第201-300条
    以此类推。咱们可以循环递增cstart的值,直到请求回来的页面没有新的出版物为止。

完整代码示例

<?php
include('simple_html_dom.php');

// 目标用户ID和基础URL
$userId = 'Sx4G9YgAAAAJ';
$baseUrl = "https://scholar.google.se/citations?user={$userId}&hl=&view_op=list_works&pagesize=100";

// 初始化存储所有出版物的数组
$allPublications = [];
$cstart = 0;
$hasMore = true;

// 模拟浏览器请求头,避免被反爬拦截
$opts = [
    'http' => [
        'header' => "User-Agent: Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36\r\n"
    ]
];
$context = stream_context_create($opts);

while ($hasMore) {
    // 构造当前页的URL
    $currentUrl = $baseUrl . "&cstart={$cstart}";
    
    // 加载页面,使用自定义上下文
    $html = file_get_html($currentUrl, false, $context);
    
    if (!$html) {
        echo "请求失败或页面无法加载,cstart: {$cstart}\n";
        break;
    }
    
    // 提取当前页的出版物(选择器需根据页面实际结构调整)
    $publications = $html->find('div.gs_r.gs_or.gs_scl');
    if (empty($publications)) {
        $hasMore = false;
        break;
    }
    
    // 处理每一条出版物,提取核心信息
    foreach ($publications as $pub) {
        $title = $pub->find('h3.gs_rt a', 0)->plaintext ?? '无标题';
        $authors = $pub->find('div.gs_a', 0)->plaintext ?? '无作者';
        $year = preg_match('/(\d{4})/', $authors, $matches) ? $matches[1] : '无年份';
        
        $allPublications[] = [
            'title' => $title,
            'authors' => $authors,
            'year' => $year
        ];
    }
    
    // 递增cstart,准备下一页
    $cstart += 100;
    
    // 关键:添加延迟,避免请求过于频繁触发反爬
    sleep(2);
    
    // 释放资源
    $html->clear();
    unset($html);
}

// 输出结果
echo "共抓取到 " . count($allPublications) . " 条出版物\n";
print_r($allPublications);
?>

重要注意事项

  • 反爬规避:一定要添加请求头模拟浏览器,并且设置合理的请求延迟(2-5秒为宜),否则很容易被Google Scholar封禁IP。如果爬取量很大,建议搭配代理IP池使用。
  • 选择器维护:Google Scholar的页面结构可能不定期更新,你需要定期检查find方法中的选择器是否仍然有效。比如如果div.gs_r.gs_or.gs_scl失效了,要根据新的页面HTML调整。
  • 异常处理:代码中加入了基础的失败判断,实际使用中可以进一步完善——比如添加HTTP错误码处理、请求重试机制等。

内容的提问来源于stack exchange,提问作者kores59

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.20 10:09:38