You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python Selenium XPath定位动态元素爬取图书馆搜索结果遇阻

问题描述

给定条件

目标图书馆搜索链接:

url = 'https://digi.kansalliskirjasto.fi/search?query=economic%20crisis&orderBy=RELEVANCE'

需求:提取该页面20条搜索结果中红框标注的有用信息。

初始代码问题

使用Selenium直接定位元素时返回空列表,代码如下:

from selenium import webdriver
from webdriver_manager.chrome import ChromeDriverManager
from selenium.webdriver.common.by import By
from selenium.webdriver.chrome.options import Options
from selenium.webdriver.chrome.service import Service

def run_selenium(URL):
    options = Options()
    options.add_argument("--remote-debugging-port=9222"),
    options.headless = True
    
    driver = webdriver.Chrome(service=Service(ChromeDriverManager().install()), options=options)
    
    driver.get(URL)
    pt = "//app-digiweb/ng-component/section/div/div/app-binding-search-results/div/div"
    medias = driver.find_elements(By.XPATH, pt) # 预期获取20个元素
    print(medias) # 实际输出:[]
    print("#"*100)
    for i, v in enumerate(medias):
        print(i, v.get_attribute("innerHTML"))

if __name__ == '__main__':
    url = 'https://digi.kansalliskirjasto.fi/search?query=economic%20crisis&orderBy=RELEVANCE'
    run_selenium(URL=url)

疑问:页面元素的id和class为动态属性,使用Selenium的driver.find_elements()是否应该返回包含全部信息的20个元素?

更新后的低效方案

改用WebDriverWait结合try-except逻辑可正常获取结果,但耗时长达15.22秒,代码如下:

import time
from bs4 import BeautifulSoup
from selenium import webdriver
from webdriver_manager.chrome import ChromeDriverManager
from selenium.webdriver.common.by import By
from selenium.webdriver.chrome.service import Service
from selenium.webdriver.chrome.options import Options
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from selenium.common import exceptions

def get_all_search_details(URL):
    st_t = time.time()
    SEARCH_RESULTS = {}
    options = Options()
    options.headless = True    
    options.add_argument("--remote-debugging-port=9222")
    options.add_argument("--no-sandbox")
    options.add_argument("--disable-gpu")
    options.add_argument("--disable-dev-shm-usage")
    options.add_argument("--disable-extensions")
    options.add_experimental_option("excludeSwitches", ["enable-automation"])
    options.add_experimental_option('useAutomationExtension', False)
    driver =webdriver.Chrome(service=Service(ChromeDriverManager().install()), options=options)
    driver.get(URL)
    print(f"Scraping {driver.current_url}")
    try:
        medias = WebDriverWait(driver,timeout=10,).until(EC.presence_of_all_elements_located((By.CLASS_NAME, 'result-row')))
        for media_idx, media_elem in enumerate(medias):
            outer_html = media_elem.get_attribute('outerHTML')
            result = scrap_newspaper(outer_html) # 自定义提取信息的函数
            SEARCH_RESULTS[f"result_{media_idx}"] = result
    except exceptions.StaleElementReferenceException as e:
        print(f"Selenium: {type(e).__name__}: {e.args}")
        return
    except exceptions.NoSuchElementException as e:
        print(f"Selenium: {type(e).__name__}: {e.args}")
        return
    except exceptions.TimeoutException as e:
        print(f"Selenium: {type(e).__name__}: {e.args}")
        return
    except exceptions.WebDriverException as e:
        print(f"Selenium: {type(e).__name__}: {e.args}")
        return
    except exceptions.SessionNotCreatedException as e:
        print(f"Selenium: {type(e).__name__}: {e.args}")
        return
    except Exception as e:
        print(f"Selenium: {type(e).__name__} line {e.__traceback__.tb_lineno} of {__file__}: {e.args}")
        return
    except:
        print(f"Selenium General Exception: {URL}")
        return
    print(f"		Found {len(medias)} media(s) => {len(SEARCH_RESULTS)} search result(s)	Elapsed_t: {time.time()-st_t:.2f} s")
    return SEARCH_RESULTS

if __name__ == '__main__':
    url = 'https://digi.kansalliskirjasto.fi'
    get_all_search_details(URL=url)

运行结果:

Found 20 media(s) => 20 search result(s) Elapsed_t: 15.22 s

寻求更高效的解决方案。


高效解决方案

1. 优化元素定位与等待策略

避免依赖动态class/id,改用稳定的标签结构+模糊匹配class定位;同时用visibility_of_all_elements_located替代presence_of_all_elements_located,确保元素可见后再操作,减少后续 stale 异常:

medias = WebDriverWait(driver, 10).until(
    EC.visibility_of_all_elements_located((By.XPATH, "//div[contains(@class, 'result-row')]"))
)

2. 批量提取HTML后解析

一次性获取结果父容器的HTML,用BeautifulSoup批量解析,减少Selenium与浏览器的交互次数(这是耗时的核心原因):

# 等待结果容器加载完成
results_container = WebDriverWait(driver, 10).until(
    EC.visibility_of_element_located((By.XPATH, "//app-binding-search-results"))
)
# 一次性获取所有结果的HTML
container_html = results_container.get_attribute('innerHTML')
soup = BeautifulSoup(container_html, 'html.parser')
# 批量提取结果行
result_rows = soup.find_all('div', class_=lambda c: c and 'result-row' in c)

for idx, row in enumerate(result_rows):
    # 直接用BeautifulSoup提取所需信息,示例:标题、来源
    title = row.find('h3').get_text(strip=True) if row.find('h3') else ''
    source = row.find('div', class_=lambda c: c and 'source' in c).get_text(strip=True) if row.find('div', class_=lambda c: c and 'source' in c) else ''
    SEARCH_RESULTS[f"result_{idx}"] = {'title': title, 'source': source}

3. 优化Chrome启动参数

移除不必要参数,添加性能优化项,减少页面加载资源:

options = Options()
options.headless = True
# 基础必要参数
options.add_argument("--no-sandbox")
options.add_argument("--disable-dev-shm-usage")
# 性能优化参数
options.add_argument("--disable-images")  # 禁用图片加载,大幅提速
options.add_argument("--blink-settings=imagesEnabled=false")
options.page_load_strategy = 'eager'  # 只等待DOM加载完成,无需等待所有资源
options.add_argument("--disable-extensions")

4. 复用浏览器实例

若需多次爬取,不要每次重新启动浏览器,复用已有的driver实例,避免重复初始化的开销。

5. 直接调用API(最优解)

打开浏览器开发者工具的Network标签,搜索关键词查看是否有返回搜索结果的XHR/fetch请求。若找到对应API接口,直接用requests调用,速度比Selenium快10倍以上:

import requests

api_url = "https://digi.kansalliskirjasto.fi/api/search"
params = {
    'query': 'economic crisis',
    'orderBy': 'RELEVANCE',
    'limit': 20
}
response = requests.get(api_url, params=params)
data = response.json()
# 解析JSON提取信息
for idx, item in enumerate(data['results']):
    SEARCH_RESULTS[f"result_{idx}"] = {
        'title': item['title'],
        'source': item['source'],
        # 其他字段根据API返回结构调整
    }

内容的提问来源于stack exchange,提问作者farid

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.15 13:30:51