You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

用Python的BeautifulSoup获取专利网页公开日期及代码差异问题解决

Espacenet专利公开日期爬取问题解决

为什么审查元素的代码和requests返回的HTML不一致?

因为Espacenet页面采用动态渲染机制:浏览器打开页面时仅加载基础静态HTML,后续通过JavaScript请求后端数据并动态生成页面内容(包括你看到的公开日期标签)。而requests库只能获取初始静态HTML,不会执行JavaScript,因此无法找到动态生成的元素。

解决方法

方法1:用Selenium模拟浏览器(直观易上手)

Selenium可以模拟真实浏览器行为,自动执行JS渲染完整页面后再提取目标元素:

  1. 先安装Selenium库及对应浏览器驱动(如ChromeDriver)
  2. 编写代码等待元素加载完成后提取内容

示例代码:

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC

# 初始化Chrome浏览器(确保ChromeDriver路径配置正确)
driver = webdriver.Chrome()
target_url = 'https://worldwide.espacenet.com/patent/search/family/054437790/publication/CN105030410A?q=CN105030410'
driver.get(target_url)

# 等待公开日期元素加载完成,超时时间10秒
wait = WebDriverWait(driver, 10)
date_element = wait.until(
    EC.presence_of_element_located((By.ID, 'biblio-publication-date-content'))
)

# 提取并打印公开日期
publication_date = date_element.text
print('公开日期:', publication_date)

# 关闭浏览器
driver.quit()

方法2:直接调用API接口(效率更高,推荐)

通过浏览器抓包可找到Espacenet加载数据的API接口,直接请求JSON数据比模拟浏览器更高效:

import requests

# 抓包获取的API接口(参数对应目标专利)
api_url = 'https://worldwide.espacenet.com/rest-services/published-data/publication/CN/105030410A/biblio'
headers = {'Accept': 'application/json'}

response = requests.get(api_url, headers=headers)
data = response.json()

# 从JSON结构中提取公开日期
publication_date = data['biblio']['publicationReference']['date']
print('公开日期:', publication_date)

注意:API接口可能随网站更新变化,若失效需重新抓包获取最新接口。

内容的提问来源于stack exchange,提问作者Umberto

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.06 04:25:18