You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何爬取Gallica网页中「informations détaillées」区域数据?

爬取Gallica「informations détaillées」数据的可行方案

一、修复Selenium爬取问题

你之前用Selenium失败大概率是没做好等待元素加载+模拟点击这两步,毕竟这个区域是点击展开的动态内容。下面是修正后的完整流程和代码:

关键操作步骤

  1. 先搞定浏览器驱动:比如用Chrome的话,下载和你Chrome版本匹配的chromedriver,把它放到Python能找到的路径里。
  2. 访问目标URL后,别直接找元素,得等页面加载完,再找到「informations détaillées」的触发按钮,模拟点击。
  3. 等展开的详情区域加载出来,再提取数据。
  4. 循环处理500个URL,最后把数据存进SQL库。

修正后的代码示例

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
import sqlite3
import time

# 初始化Chrome浏览器,可选无头模式(后台运行)
options = webdriver.ChromeOptions()
# options.add_argument('--headless=new')  # 放开注释就后台运行
driver = webdriver.Chrome(options=options)
wait = WebDriverWait(driver, 15)  # 设置15秒超时,应对页面加载慢的情况

# 替换成你的500个URL列表
url_list = ["https://gallica.bnf.fr/ark:/12148/bpt6k12345678"]

# 连接SQLite数据库(换成MySQL/PostgreSQL的话改下连接代码就行)
conn = sqlite3.connect('gallica_data.db')
cursor = conn.cursor()
# 创建表,字段根据你要爬的内容调整
cursor.execute('''CREATE TABLE IF NOT EXISTS gallica_details
                  (url TEXT PRIMARY KEY, title TEXT, author TEXT, publication_info TEXT, full_details TEXT)''')

for idx, url in enumerate(url_list):
    try:
        print(f"正在处理第{idx+1}个URL:{url}")
        driver.get(url)
        
        # 定位并点击「informations détaillées」按钮(用XPATH定位,可根据实际页面调整)
        # 先等按钮可点击再操作,避免元素没加载完报错
        detail_btn = wait.until(EC.element_to_be_clickable(
            (By.XPATH, "//*[contains(text(), 'informations détaillées')]")
        ))
        detail_btn.click()
        time.sleep(1)  # 加个短等待,确保展开区域加载
        
        # 等待详情区域出现,提取数据
        detail_section = wait.until(EC.presence_of_element_located(
            (By.CLASS_NAME, "notice-detailles")  # 这个class是Gallica实际用的,可验证
        ))
        
        # 提取具体字段,根据页面实际结构调整
        title = driver.find_element(By.CLASS_NAME, "notice-title").text
        author = driver.find_element(By.CLASS_NAME, "notice-auteur").text if driver.find_elements(By.CLASS_NAME, "notice-auteur") else "无作者信息"
        pub_info = driver.find_element(By.CLASS_NAME, "notice-publication").text if driver.find_elements(By.CLASS_NAME, "notice-publication") else "无出版信息"
        full_details = detail_section.text
        
        # 插入数据库,避免重复数据用INSERT OR REPLACE
        cursor.execute("INSERT OR REPLACE INTO gallica_details VALUES (?, ?, ?, ?, ?)",
                      (url, title, author, pub_info, full_details))
        conn.commit()
        print(f"✅ 成功保存:{url}")
        
    except Exception as e:
        print(f"❌ 处理{url}出错:{str(e)}")
        continue

# 收尾工作
driver.quit()
conn.close()
print("所有URL处理完成!")

常见错误排查

  • 驱动版本不匹配:去Chrome设置里看版本,下载对应版本的chromedriver,别瞎下最新的。
  • 元素定位不到:按F12打开开发者工具,找到按钮的真实HTML标签(可能是<button>或<a>),调整XPATH或CSS选择器。比如按钮的class是btn-details,就用By.CLASS_NAME, "btn-details"。
  • 反爬拦截:Gallica偶尔会查人机,别爬太快,加time.sleep(2)间隔,或者用无头模式时加个user-agent:options.add_argument('user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36')

二、更高效的方案:直接调用Gallica API

Selenium跑起来慢,批量爬500个URL太费时间。其实Gallica有官方API,能直接拿详情数据,不用模拟点击:

操作步骤

  1. 打开Gallica的某个文档页面,按F12切到「网络」标签,点击「informations détaillées」,观察新出现的请求,找到类似https://gallica.bnf.fr/services/engine/search/getDetail?ark=xxx的API接口。
  2. 这个API的参数就是文档的ARK ID(比如URL里的bpt6k12345678),直接用requests请求就行。

API爬取代码示例

import requests
import sqlite3
import time

conn = sqlite3.connect('gallica_data.db')
cursor = conn.cursor()
cursor.execute('''CREATE TABLE IF NOT EXISTS gallica_api_details
                  (url TEXT PRIMARY KEY, ark_id TEXT, detail_json TEXT)''')

url_list = ["https://gallica.bnf.fr/ark:/12148/bpt6k12345678"]

for idx, url in enumerate(url_list):
    try:
        print(f"正在处理第{idx+1}个URL:{url}")
        # 从URL里提取ARK ID
        ark_id = url.split('/')[-1]
        # 官方API地址,参数就是ark_id
        api_url = f"https://gallica.bnf.fr/services/engine/search/getDetail?ark={ark_id}"
        
        # 加请求头,避免被拦截
        headers = {
            'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36'
        }
        response = requests.get(api_url, headers=headers)
        response.raise_for_status()  # 检查请求是否成功
        detail_data = response.json()
        
        # 存数据库,也可以把JSON里的字段拆出来单独存
        cursor.execute("INSERT OR REPLACE INTO gallica_api_details VALUES (?, ?, ?)",
                      (url, ark_id, str(detail_data)))
        conn.commit()
        print(f"✅ 成功获取数据:{url}")
        time.sleep(1)  # 控制请求频率,别给服务器压力
        
    except Exception as e:
        print(f"❌ 处理{url}出错:{str(e)}")
        continue

conn.close()
print("所有URL处理完成!")

这个方法比Selenium快好几倍,而且资源占用少,适合批量爬取。

三、SQL存储注意事项

  • 提前设计好表结构,比如要存作者、标题、出版日期这些字段,就把对应字段加上,别全存成文本。
  • 批量插入时用事务,或者每10个URL提交一次,提高效率。
  • 处理重复URL,用INSERT OR REPLACE或者先查询是否存在再插入,避免数据库里有重复数据。

内容的提问来源于stack exchange,提问作者Moran Hanane

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.14 13:32:16