使用Selenium爬取Google专利搜索结果时缺漏链接如何解决
Selenium爬取Google专利结果缺漏链接问题解决思路
问题根源
漏抓的核心原因是Google专利的搜索结果存在两种结构:
- 带PDF预览的条目:
search-result-item外层元素直接绑定href属性,指向PDF文件 - 非PDF类条目:
href属性绑定在search-result-item的子级<a>标签上,之前直接取外层元素的href就会拿到空值,直接被过滤掉了
另外当前代码没有加等待逻辑,可能元素还没完全加载完成就执行了抓取,也会导致漏抓。
解决步骤
- 新增显式等待逻辑,确保页面所有搜索结果元素渲染完成后再执行抓取
- 修改链接提取逻辑,统一从每个搜索结果的子级
<a>标签提取href,覆盖两种条目结构 - 若需爬取全量结果,还要补充滚动加载逻辑,Google专利默认仅加载首屏内容,滚动到页面底部才会加载下一批结果
修正后可运行代码
import selenium from selenium import webdriver from selenium.webdriver.firefox.options import Options as options from selenium.webdriver.firefox.service import Service from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC new_driver_path = r"C:/Users/alexe/Desktop/Apple/PatentSearch/geckodriver-v0.30.0-win64/geckodriver.exe" ops = options() serv = Service(new_driver_path) browser1 = selenium.webdriver.Firefox(service=serv, options=ops) browser1.get("https://patents.google.com/?assignee=apple&after=priority:20150101&sort=new") # 显式等待最多10秒,直到所有搜索结果元素加载完成 wait = WebDriverWait(browser1, 10) elements = wait.until(EC.presence_of_all_elements_located((By.CLASS_NAME, "search-result-item"))) links = [] for elem in elements: # 从子级a标签提取href,覆盖所有类型条目 a_tag = elem.find_element(By.TAG_NAME, "a") href = a_tag.get_attribute('href') if href: links.append(href) links = set(links) for href in links: print(href)
内容的提问来源于stack exchange,提问作者Alexey Lantsov
相关产品推荐
相关产品推荐

