使用SimpleHtmlDom爬取Google Scholar个人主页文献的数量限制问题
解决Google Scholar个人出版物全量爬取问题
我懂你现在的困扰——用SimpleHtmlDom只能抓到当前页面显示的内容,哪怕把pagesize调到100,超过这个数量的出版物还是拿不到。其实核心问题是Google Scholar的内容是分页加载的,咱们得通过循环遍历所有分页来获取全部内容。
核心思路
Google Scholar的分页是通过cstart参数控制的:
cstart=0:对应第1-100条(搭配pagesize=100)cstart=100:对应第101-200条cstart=200:对应第201-300条
以此类推。咱们可以循环递增cstart的值,直到请求回来的页面没有新的出版物为止。
完整代码示例
<?php include('simple_html_dom.php'); // 目标用户ID和基础URL $userId = 'Sx4G9YgAAAAJ'; $baseUrl = "https://scholar.google.se/citations?user={$userId}&hl=&view_op=list_works&pagesize=100"; // 初始化存储所有出版物的数组 $allPublications = []; $cstart = 0; $hasMore = true; // 模拟浏览器请求头,避免被反爬拦截 $opts = [ 'http' => [ 'header' => "User-Agent: Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36\r\n" ] ]; $context = stream_context_create($opts); while ($hasMore) { // 构造当前页的URL $currentUrl = $baseUrl . "&cstart={$cstart}"; // 加载页面,使用自定义上下文 $html = file_get_html($currentUrl, false, $context); if (!$html) { echo "请求失败或页面无法加载,cstart: {$cstart}\n"; break; } // 提取当前页的出版物(选择器需根据页面实际结构调整) $publications = $html->find('div.gs_r.gs_or.gs_scl'); if (empty($publications)) { $hasMore = false; break; } // 处理每一条出版物,提取核心信息 foreach ($publications as $pub) { $title = $pub->find('h3.gs_rt a', 0)->plaintext ?? '无标题'; $authors = $pub->find('div.gs_a', 0)->plaintext ?? '无作者'; $year = preg_match('/(\d{4})/', $authors, $matches) ? $matches[1] : '无年份'; $allPublications[] = [ 'title' => $title, 'authors' => $authors, 'year' => $year ]; } // 递增cstart,准备下一页 $cstart += 100; // 关键:添加延迟,避免请求过于频繁触发反爬 sleep(2); // 释放资源 $html->clear(); unset($html); } // 输出结果 echo "共抓取到 " . count($allPublications) . " 条出版物\n"; print_r($allPublications); ?>
重要注意事项
- 反爬规避:一定要添加请求头模拟浏览器,并且设置合理的请求延迟(2-5秒为宜),否则很容易被Google Scholar封禁IP。如果爬取量很大,建议搭配代理IP池使用。
- 选择器维护:Google Scholar的页面结构可能不定期更新,你需要定期检查
find方法中的选择器是否仍然有效。比如如果div.gs_r.gs_or.gs_scl失效了,要根据新的页面HTML调整。 - 异常处理:代码中加入了基础的失败判断,实际使用中可以进一步完善——比如添加HTTP错误码处理、请求重试机制等。
内容的提问来源于stack exchange,提问作者kores59
相关产品推荐
相关产品推荐

