Selenium XPath匹配问题:如何仅获取'Games'板块数据排除'Virtual Games'
解决Selenium爬虫XPath误匹配“Virtual Games”的问题
问题背景
目标页面:https://www.askgamblers.com/online-casinos/reviews/yukon-gold-casino-casino
需求:提取Games板块下的文本内容
问题:原代码中XPath使用contains(.,'Games'),会同时匹配到“Virtual Games”板块的内容,导致获取的数据不符合预期。
关键解决方案
原XPath的核心问题:
contains(.,'Games')会匹配所有包含“Games”的文本,包括“Virtual Games”- 使用绝对路径
//会全局查找元素,而非限定在当前section范围内
修改后的XPath需满足:
- 精确匹配h2文本为“Games”(用
normalize-space处理可能的空格问题) - 仅获取该h2之后、下一个同级h2(如“Virtual Games”)之前的p标签内容
- 使用相对路径
.//限定在当前section内查找
修改后的核心XPath示例
.//h2[normalize-space(text())='Games']/following-sibling::p[following-sibling::h2[normalize-space(text())='Virtual Games']]
修改后的完整代码
from selenium import webdriver import time from selenium.webdriver.chrome.service import Service from selenium.webdriver.common.by import By from webdriver_manager.chrome import ChromeDriverManager from selenium.webdriver.support.wait import WebDriverWait options = webdriver.ChromeOptions() options.add_argument("--no-sandbox") options.add_argument("--disable-gpu") options.add_argument("--window-size=1920x1080") options.add_argument("--disable-extensions") driver = webdriver.Chrome(service=Service(ChromeDriverManager().install())) wait = WebDriverWait(driver, 20) URL = 'https://www.askgamblers.com/online-casinos/reviews/yukon-gold-casino-casino' driver.get(URL) # 等待页面核心区块加载完成 wait.until(lambda d: d.find_element(By.XPATH, "//section[@class='review-text richtext']")) data = driver.find_elements(By.XPATH,"//section[@class='review-text richtext']") for row in data: try: # 获取Games板块下的所有段落文本 game_paragraphs = row.find_elements(By.XPATH,".//h2[normalize-space(text())='Games']/following-sibling::p[following-sibling::h2[normalize-space(text())='Virtual Games']]") para0 = '\n'.join([p.text for p in game_paragraphs]) print(para0) except Exception as e: print(f"获取数据失败: {e}") pass driver.quit()
改动说明
- 增加页面加载等待逻辑,确保元素渲染完成
- 改用相对路径
.//,限定在当前section范围内查找元素 - 用
normalize-space(text())='Games'精确匹配标题,彻底排除“Virtual Games”的误匹配 - 用
following-sibling::h2[normalize-space(text())='Virtual Games']限定内容范围,只取“Games”到“Virtual Games”之间的段落 - 批量获取所有相关p标签并拼接成完整文本,提升内容完整性
内容的提问来源于stack exchange,提问作者Amen Aziz
相关产品推荐
相关产品推荐

