Python Selenium XPath定位动态元素爬取图书馆搜索结果遇阻
问题描述
给定条件
目标图书馆搜索链接:
url = 'https://digi.kansalliskirjasto.fi/search?query=economic%20crisis&orderBy=RELEVANCE'
需求:提取该页面20条搜索结果中红框标注的有用信息。
初始代码问题
使用Selenium直接定位元素时返回空列表,代码如下:
from selenium import webdriver from webdriver_manager.chrome import ChromeDriverManager from selenium.webdriver.common.by import By from selenium.webdriver.chrome.options import Options from selenium.webdriver.chrome.service import Service def run_selenium(URL): options = Options() options.add_argument("--remote-debugging-port=9222"), options.headless = True driver = webdriver.Chrome(service=Service(ChromeDriverManager().install()), options=options) driver.get(URL) pt = "//app-digiweb/ng-component/section/div/div/app-binding-search-results/div/div" medias = driver.find_elements(By.XPATH, pt) # 预期获取20个元素 print(medias) # 实际输出:[] print("#"*100) for i, v in enumerate(medias): print(i, v.get_attribute("innerHTML")) if __name__ == '__main__': url = 'https://digi.kansalliskirjasto.fi/search?query=economic%20crisis&orderBy=RELEVANCE' run_selenium(URL=url)
疑问:页面元素的id和class为动态属性,使用Selenium的driver.find_elements()是否应该返回包含全部信息的20个元素?
更新后的低效方案
改用WebDriverWait结合try-except逻辑可正常获取结果,但耗时长达15.22秒,代码如下:
import time from bs4 import BeautifulSoup from selenium import webdriver from webdriver_manager.chrome import ChromeDriverManager from selenium.webdriver.common.by import By from selenium.webdriver.chrome.service import Service from selenium.webdriver.chrome.options import Options from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from selenium.common import exceptions def get_all_search_details(URL): st_t = time.time() SEARCH_RESULTS = {} options = Options() options.headless = True options.add_argument("--remote-debugging-port=9222") options.add_argument("--no-sandbox") options.add_argument("--disable-gpu") options.add_argument("--disable-dev-shm-usage") options.add_argument("--disable-extensions") options.add_experimental_option("excludeSwitches", ["enable-automation"]) options.add_experimental_option('useAutomationExtension', False) driver =webdriver.Chrome(service=Service(ChromeDriverManager().install()), options=options) driver.get(URL) print(f"Scraping {driver.current_url}") try: medias = WebDriverWait(driver,timeout=10,).until(EC.presence_of_all_elements_located((By.CLASS_NAME, 'result-row'))) for media_idx, media_elem in enumerate(medias): outer_html = media_elem.get_attribute('outerHTML') result = scrap_newspaper(outer_html) # 自定义提取信息的函数 SEARCH_RESULTS[f"result_{media_idx}"] = result except exceptions.StaleElementReferenceException as e: print(f"Selenium: {type(e).__name__}: {e.args}") return except exceptions.NoSuchElementException as e: print(f"Selenium: {type(e).__name__}: {e.args}") return except exceptions.TimeoutException as e: print(f"Selenium: {type(e).__name__}: {e.args}") return except exceptions.WebDriverException as e: print(f"Selenium: {type(e).__name__}: {e.args}") return except exceptions.SessionNotCreatedException as e: print(f"Selenium: {type(e).__name__}: {e.args}") return except Exception as e: print(f"Selenium: {type(e).__name__} line {e.__traceback__.tb_lineno} of {__file__}: {e.args}") return except: print(f"Selenium General Exception: {URL}") return print(f" Found {len(medias)} media(s) => {len(SEARCH_RESULTS)} search result(s) Elapsed_t: {time.time()-st_t:.2f} s") return SEARCH_RESULTS if __name__ == '__main__': url = 'https://digi.kansalliskirjasto.fi' get_all_search_details(URL=url)
运行结果:
Found 20 media(s) => 20 search result(s) Elapsed_t: 15.22 s
寻求更高效的解决方案。
高效解决方案
1. 优化元素定位与等待策略
避免依赖动态class/id,改用稳定的标签结构+模糊匹配class定位;同时用visibility_of_all_elements_located替代presence_of_all_elements_located,确保元素可见后再操作,减少后续 stale 异常:
medias = WebDriverWait(driver, 10).until( EC.visibility_of_all_elements_located((By.XPATH, "//div[contains(@class, 'result-row')]")) )
2. 批量提取HTML后解析
一次性获取结果父容器的HTML,用BeautifulSoup批量解析,减少Selenium与浏览器的交互次数(这是耗时的核心原因):
# 等待结果容器加载完成 results_container = WebDriverWait(driver, 10).until( EC.visibility_of_element_located((By.XPATH, "//app-binding-search-results")) ) # 一次性获取所有结果的HTML container_html = results_container.get_attribute('innerHTML') soup = BeautifulSoup(container_html, 'html.parser') # 批量提取结果行 result_rows = soup.find_all('div', class_=lambda c: c and 'result-row' in c) for idx, row in enumerate(result_rows): # 直接用BeautifulSoup提取所需信息,示例:标题、来源 title = row.find('h3').get_text(strip=True) if row.find('h3') else '' source = row.find('div', class_=lambda c: c and 'source' in c).get_text(strip=True) if row.find('div', class_=lambda c: c and 'source' in c) else '' SEARCH_RESULTS[f"result_{idx}"] = {'title': title, 'source': source}
3. 优化Chrome启动参数
移除不必要参数,添加性能优化项,减少页面加载资源:
options = Options() options.headless = True # 基础必要参数 options.add_argument("--no-sandbox") options.add_argument("--disable-dev-shm-usage") # 性能优化参数 options.add_argument("--disable-images") # 禁用图片加载,大幅提速 options.add_argument("--blink-settings=imagesEnabled=false") options.page_load_strategy = 'eager' # 只等待DOM加载完成,无需等待所有资源 options.add_argument("--disable-extensions")
4. 复用浏览器实例
若需多次爬取,不要每次重新启动浏览器,复用已有的driver实例,避免重复初始化的开销。
5. 直接调用API(最优解)
打开浏览器开发者工具的Network标签,搜索关键词查看是否有返回搜索结果的XHR/fetch请求。若找到对应API接口,直接用requests调用,速度比Selenium快10倍以上:
import requests api_url = "https://digi.kansalliskirjasto.fi/api/search" params = { 'query': 'economic crisis', 'orderBy': 'RELEVANCE', 'limit': 20 } response = requests.get(api_url, params=params) data = response.json() # 解析JSON提取信息 for idx, item in enumerate(data['results']): SEARCH_RESULTS[f"result_{idx}"] = { 'title': item['title'], 'source': item['source'], # 其他字段根据API返回结构调整 }
内容的提问来源于stack exchange,提问作者farid
相关产品推荐
相关产品推荐

