如何爬取Gallica网页中「informations détaillées」区域数据?
爬取Gallica「informations détaillées」数据的可行方案
一、修复Selenium爬取问题
你之前用Selenium失败大概率是没做好等待元素加载+模拟点击这两步,毕竟这个区域是点击展开的动态内容。下面是修正后的完整流程和代码:
关键操作步骤
- 先搞定浏览器驱动:比如用Chrome的话,下载和你Chrome版本匹配的
chromedriver,把它放到Python能找到的路径里。 - 访问目标URL后,别直接找元素,得等页面加载完,再找到「informations détaillées」的触发按钮,模拟点击。
- 等展开的详情区域加载出来,再提取数据。
- 循环处理500个URL,最后把数据存进SQL库。
修正后的代码示例
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC import sqlite3 import time # 初始化Chrome浏览器,可选无头模式(后台运行) options = webdriver.ChromeOptions() # options.add_argument('--headless=new') # 放开注释就后台运行 driver = webdriver.Chrome(options=options) wait = WebDriverWait(driver, 15) # 设置15秒超时,应对页面加载慢的情况 # 替换成你的500个URL列表 url_list = ["https://gallica.bnf.fr/ark:/12148/bpt6k12345678"] # 连接SQLite数据库(换成MySQL/PostgreSQL的话改下连接代码就行) conn = sqlite3.connect('gallica_data.db') cursor = conn.cursor() # 创建表,字段根据你要爬的内容调整 cursor.execute('''CREATE TABLE IF NOT EXISTS gallica_details (url TEXT PRIMARY KEY, title TEXT, author TEXT, publication_info TEXT, full_details TEXT)''') for idx, url in enumerate(url_list): try: print(f"正在处理第{idx+1}个URL:{url}") driver.get(url) # 定位并点击「informations détaillées」按钮(用XPATH定位,可根据实际页面调整) # 先等按钮可点击再操作,避免元素没加载完报错 detail_btn = wait.until(EC.element_to_be_clickable( (By.XPATH, "//*[contains(text(), 'informations détaillées')]") )) detail_btn.click() time.sleep(1) # 加个短等待,确保展开区域加载 # 等待详情区域出现,提取数据 detail_section = wait.until(EC.presence_of_element_located( (By.CLASS_NAME, "notice-detailles") # 这个class是Gallica实际用的,可验证 )) # 提取具体字段,根据页面实际结构调整 title = driver.find_element(By.CLASS_NAME, "notice-title").text author = driver.find_element(By.CLASS_NAME, "notice-auteur").text if driver.find_elements(By.CLASS_NAME, "notice-auteur") else "无作者信息" pub_info = driver.find_element(By.CLASS_NAME, "notice-publication").text if driver.find_elements(By.CLASS_NAME, "notice-publication") else "无出版信息" full_details = detail_section.text # 插入数据库,避免重复数据用INSERT OR REPLACE cursor.execute("INSERT OR REPLACE INTO gallica_details VALUES (?, ?, ?, ?, ?)", (url, title, author, pub_info, full_details)) conn.commit() print(f"✅ 成功保存:{url}") except Exception as e: print(f"❌ 处理{url}出错:{str(e)}") continue # 收尾工作 driver.quit() conn.close() print("所有URL处理完成!")
常见错误排查
- 驱动版本不匹配:去Chrome设置里看版本,下载对应版本的chromedriver,别瞎下最新的。
- 元素定位不到:按F12打开开发者工具,找到按钮的真实HTML标签(可能是
<button>或<a>),调整XPATH或CSS选择器。比如按钮的class是btn-details,就用By.CLASS_NAME, "btn-details"。 - 反爬拦截:Gallica偶尔会查人机,别爬太快,加
time.sleep(2)间隔,或者用无头模式时加个user-agent:options.add_argument('user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36')
二、更高效的方案:直接调用Gallica API
Selenium跑起来慢,批量爬500个URL太费时间。其实Gallica有官方API,能直接拿详情数据,不用模拟点击:
操作步骤
- 打开Gallica的某个文档页面,按F12切到「网络」标签,点击「informations détaillées」,观察新出现的请求,找到类似
https://gallica.bnf.fr/services/engine/search/getDetail?ark=xxx的API接口。 - 这个API的参数就是文档的ARK ID(比如URL里的
bpt6k12345678),直接用requests请求就行。
API爬取代码示例
import requests import sqlite3 import time conn = sqlite3.connect('gallica_data.db') cursor = conn.cursor() cursor.execute('''CREATE TABLE IF NOT EXISTS gallica_api_details (url TEXT PRIMARY KEY, ark_id TEXT, detail_json TEXT)''') url_list = ["https://gallica.bnf.fr/ark:/12148/bpt6k12345678"] for idx, url in enumerate(url_list): try: print(f"正在处理第{idx+1}个URL:{url}") # 从URL里提取ARK ID ark_id = url.split('/')[-1] # 官方API地址,参数就是ark_id api_url = f"https://gallica.bnf.fr/services/engine/search/getDetail?ark={ark_id}" # 加请求头,避免被拦截 headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36' } response = requests.get(api_url, headers=headers) response.raise_for_status() # 检查请求是否成功 detail_data = response.json() # 存数据库,也可以把JSON里的字段拆出来单独存 cursor.execute("INSERT OR REPLACE INTO gallica_api_details VALUES (?, ?, ?)", (url, ark_id, str(detail_data))) conn.commit() print(f"✅ 成功获取数据:{url}") time.sleep(1) # 控制请求频率,别给服务器压力 except Exception as e: print(f"❌ 处理{url}出错:{str(e)}") continue conn.close() print("所有URL处理完成!")
这个方法比Selenium快好几倍,而且资源占用少,适合批量爬取。
三、SQL存储注意事项
- 提前设计好表结构,比如要存作者、标题、出版日期这些字段,就把对应字段加上,别全存成文本。
- 批量插入时用事务,或者每10个URL提交一次,提高效率。
- 处理重复URL,用
INSERT OR REPLACE或者先查询是否存在再插入,避免数据库里有重复数据。
内容的提问来源于stack exchange,提问作者Moran Hanane
相关产品推荐
相关产品推荐

