You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

REU项目Python爬取专利文本:如何避免Serpapi单专利消耗搜索额度?

解决专利全文爬取的API额度消耗问题
  • 先核查Serpapi搜索接口的返回字段:你当前使用的Google Patents搜索API,返回结果里可能已经包含了专利的核心文本内容(比如摘要、说明书片段甚至完整全文),无需额外调用Details API。去查看Serpapi的响应结构文档,找类似description、full_description或claims这类字段,要是能直接拿到所需内容,就能省去每个专利单独消耗的搜索额度。

  • 改用公开网页爬取(需遵守合规规则):如果API额度紧张,可以直接爬取Google Patents的公开网页内容。用Python的requests配合BeautifulSoup库,先请求搜索结果页面解析出所有专利详情链接,再逐个访问详情页提取文本。注意给请求加2-3秒的延时,避免触发反爬机制,同时严格遵守Google的robots协议。示例代码如下:

import requests
from bs4 import BeautifulSoup
import time

# 替换成你的查询关键词
search_query = "你的特定查询关键词"
search_url = f"https://patents.google.com/?q={search_query}"
headers = {"User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36"}

# 获取搜索结果中的专利详情链接
response = requests.get(search_url, headers=headers)
soup = BeautifulSoup(response.text, "html.parser")
patent_links = [a["href"] for a in soup.select("a.result-item")]

# 遍历链接爬取全文并保存
for link in patent_links:
    full_link = f"https://patents.google.com{link}"
    detail_resp = requests.get(full_link, headers=headers)
    detail_soup = BeautifulSoup(detail_resp.text, "html.parser")
    
    # 提取专利文本(根据页面结构调整选择器)
    patent_text = detail_soup.select_one("div.patent-text").get_text(strip=True)
    
    # 按专利ID命名文件并保存
    patent_id = link.split("/")[-1]
    with open(f"{patent_id}.txt", "w", encoding="utf-8") as f:
        f.write(patent_text)
    
    time.sleep(2)  # 延时避免反爬
  • 利用批量专利数据源:像USPTO(美国专利商标局)提供了批量专利数据下载服务,你可以在其数据库中筛选符合特定查询条件的专利,直接下载包含全文的压缩包,解压后即可提取所需文本,完全无需消耗API额度。

内容的提问来源于stack exchange,提问作者wanxiang

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.22 08:12:33