You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Selenium爬取Google专利搜索结果时缺漏链接如何解决

Selenium爬取Google专利结果缺漏链接问题解决思路

问题根源

漏抓的核心原因是Google专利的搜索结果存在两种结构:

  • 带PDF预览的条目:search-result-item 外层元素直接绑定href属性,指向PDF文件
  • 非PDF类条目:href属性绑定在search-result-item的子级<a>标签上,之前直接取外层元素的href就会拿到空值,直接被过滤掉了
    另外当前代码没有加等待逻辑,可能元素还没完全加载完成就执行了抓取,也会导致漏抓。

解决步骤

  • 新增显式等待逻辑,确保页面所有搜索结果元素渲染完成后再执行抓取
  • 修改链接提取逻辑,统一从每个搜索结果的子级<a>标签提取href,覆盖两种条目结构
  • 若需爬取全量结果,还要补充滚动加载逻辑,Google专利默认仅加载首屏内容,滚动到页面底部才会加载下一批结果

修正后可运行代码

import selenium
from selenium import webdriver
from selenium.webdriver.firefox.options import Options as options
from selenium.webdriver.firefox.service import Service
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC

new_driver_path = r"C:/Users/alexe/Desktop/Apple/PatentSearch/geckodriver-v0.30.0-win64/geckodriver.exe"

ops = options()
serv = Service(new_driver_path)
browser1 = selenium.webdriver.Firefox(service=serv, options=ops)
browser1.get("https://patents.google.com/?assignee=apple&after=priority:20150101&sort=new")

# 显式等待最多10秒,直到所有搜索结果元素加载完成
wait = WebDriverWait(browser1, 10)
elements = wait.until(EC.presence_of_all_elements_located((By.CLASS_NAME, "search-result-item")))

links = []
for elem in elements:
    # 从子级a标签提取href,覆盖所有类型条目
    a_tag = elem.find_element(By.TAG_NAME, "a")
    href = a_tag.get_attribute('href')
    if href:
        links.append(href)

links = set(links)
for href in links:
    print(href)

内容的提问来源于stack exchange,提问作者Alexey Lantsov

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.26 16:54:02