You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

求助:使用Selenium+BeautifulSoup提取指定段落返回None的问题

问题分析与修复方案

核心问题

  1. 选择器失效:h2:-soup-contains('Virtual Games')的匹配逻辑不够精准,部分页面的“Virtual Games”标题可能存在文本格式差异,且nextSibling容易抓取到空白文本节点而非目标段落。
  2. 未等待内容加载:点击“read more”后直接解析页面,可能完整内容还未渲染完成,导致BeautifulSoup无法抓取有效数据。
  3. 冗余代码:重复导入Service、By模块,代码结构可优化。

修复步骤

  • 替换选择器:用更兼容的文本匹配方式定位标题,再精准定位其下方的段落元素。
  • 完善等待逻辑:点击“read more”后,等待内容完全展开再解析页面。
  • 精简冗余代码,细化错误提示。

修改后的完整代码

from selenium import webdriver
import time
from selenium.webdriver.chrome.service import Service
from selenium.webdriver.common.by import By
from webdriver_manager.chrome import ChromeDriverManager
from bs4 import BeautifulSoup
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.support.wait import WebDriverWait

options = webdriver.ChromeOptions()
options.add_argument("--no-sandbox")
options.add_argument("--disable-gpu")
options.add_argument("--window-size=1920x1080")
options.add_argument("--disable-extensions")
driver = webdriver.Chrome(service=Service(ChromeDriverManager().install()))

URL = 'https://www.askgamblers.com/online-casinos/countries/uk'
driver.get(URL)
time.sleep(2)
urls = []
page_links = driver.find_elements(By.XPATH, "//div[@class='card__desc']//a[starts-with(@href, '/online')]")
for link in page_links:
    href = link.get_attribute("href")
    urls.append(href)

for url in urls:
    driver.get(url)
    # 等待页面主体内容加载
    WebDriverWait(driver, 10).until(EC.presence_of_element_located((By.CLASS_NAME, "review-main")))
    
    try:
        # 等待"read more"按钮可点击并触发点击
        review_btn = WebDriverWait(driver, 20).until(EC.element_to_be_clickable((By.XPATH, "//a[@class='review-main__show']")))
        review_btn.click()
        # 等待按钮消失,确认内容已展开
        WebDriverWait(driver, 10).until(EC.invisibility_of_element_located((By.XPATH, "//a[@class='review-main__show']")))
    except:
        pass
    
    soup = BeautifulSoup(driver.page_source, "lxml")
    
    try:
        # 定位"Virtual Games"标题(兼容文本空格、大小写差异)
        virtual_games_heading = soup.find("h2", string=lambda text: text and "Virtual Games" in text.strip())
        if virtual_games_heading:
            # 优先找标题后的p标签段落
            paragraph = virtual_games_heading.find_next_sibling("p")
            if paragraph:
                print(paragraph.get_text(strip=True))
            else:
                # 若没有p标签,尝试匹配内容容器div
                paragraph = virtual_games_heading.find_next_sibling("div", class_="review-main__content")
                print(paragraph.get_text(strip=True) if paragraph else "No content found")
        else:
            print("Virtual Games section not found")
    except Exception as e:
        print(f"Error: {str(e)}")
        pass

driver.quit()

关键改动说明

  1. 选择器优化:用find("h2", string=lambda text: text and "Virtual Games" in text.strip())替代原有的soup-contains,兼容性更强,能匹配标题文本存在空格或格式差异的情况。
  2. 精准定位段落:用find_next_sibling替代nextSibling,自动跳过空白文本节点,直接定位下一个有效内容元素(优先p标签,再 fallback 到内容容器div)。
  3. 等待逻辑完善:点击“read more”后等待按钮消失,确保内容完全展开后再解析页面,避免抓取到未加载的内容。
  4. 错误提示细化:增加标题不存在、段落不存在的分支判断,输出更明确的状态信息,便于排查问题。

内容的提问来源于stack exchange,提问作者developer

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.10 04:31:12