用Python的BeautifulSoup获取专利网页公开日期及代码差异问题解决
Espacenet专利公开日期爬取问题解决
为什么审查元素的代码和requests返回的HTML不一致?
因为Espacenet页面采用动态渲染机制:浏览器打开页面时仅加载基础静态HTML,后续通过JavaScript请求后端数据并动态生成页面内容(包括你看到的公开日期标签)。而requests库只能获取初始静态HTML,不会执行JavaScript,因此无法找到动态生成的元素。
解决方法
方法1:用Selenium模拟浏览器(直观易上手)
Selenium可以模拟真实浏览器行为,自动执行JS渲染完整页面后再提取目标元素:
- 先安装Selenium库及对应浏览器驱动(如ChromeDriver)
- 编写代码等待元素加载完成后提取内容
示例代码:
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC # 初始化Chrome浏览器(确保ChromeDriver路径配置正确) driver = webdriver.Chrome() target_url = 'https://worldwide.espacenet.com/patent/search/family/054437790/publication/CN105030410A?q=CN105030410' driver.get(target_url) # 等待公开日期元素加载完成,超时时间10秒 wait = WebDriverWait(driver, 10) date_element = wait.until( EC.presence_of_element_located((By.ID, 'biblio-publication-date-content')) ) # 提取并打印公开日期 publication_date = date_element.text print('公开日期:', publication_date) # 关闭浏览器 driver.quit()
方法2:直接调用API接口(效率更高,推荐)
通过浏览器抓包可找到Espacenet加载数据的API接口,直接请求JSON数据比模拟浏览器更高效:
import requests # 抓包获取的API接口(参数对应目标专利) api_url = 'https://worldwide.espacenet.com/rest-services/published-data/publication/CN/105030410A/biblio' headers = {'Accept': 'application/json'} response = requests.get(api_url, headers=headers) data = response.json() # 从JSON结构中提取公开日期 publication_date = data['biblio']['publicationReference']['date'] print('公开日期:', publication_date)
注意:API接口可能随网站更新变化,若失效需重新抓包获取最新接口。
内容的提问来源于stack exchange,提问作者Umberto
相关产品推荐
相关产品推荐

