如何使用Selenium遍历折叠面板并抓取治疗的标题、项目与价格
解决Selenium遍历折叠面板抓取牙科价格的问题
原代码存在的问题
- 折叠面板展开逻辑有问题:用
material-icons作为选择器会匹配到页面上其他无关图标,导致部分面板无法正确展开 - 内容提取逻辑错误:遍历
col-12区块时,直接调用driver.find_element只会获取页面中第一个accordion-content的内容,而非当前遍历区块下的内容 - 未做结构化提取:只是把整块文本存入列表,没有拆分出治疗标题、项目名称和价格
修正后的实现代码
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.chrome.service import Service from selenium.common.exceptions import NoSuchElementException, ElementClickInterceptedException import time service_obj = Service(r"C:\chromedriver.exe") options = webdriver.ChromeOptions() options.add_experimental_option("detach", True) driver = webdriver.Chrome(service=service_obj, options=options) url = 'https://www.23dental.com/plans-fees/dental-fees' driver.get(url) driver.implicitly_wait(5) # 拒绝Cookie try: reject_cookies_btn = driver.find_element(By.ID, "onetrust-reject-all-handler") reject_cookies_btn.click() time.sleep(1) except NoSuchElementException: pass # 展开所有折叠面板 # 精准定位每个面板的展开按钮(每个accordion-header下的触发元素) accordion_headers = driver.find_elements(By.CLASS_NAME, 'accordion-header') for header in accordion_headers: try: # 点击展开按钮 expand_btn = header.find_element(By.CSS_SELECTOR, '.material-icons') if expand_btn.text == 'expand_more': expand_btn.click() time.sleep(0.5) except ElementClickInterceptedException: # 处理可能的遮挡问题 driver.execute_script("arguments[0].click();", expand_btn) time.sleep(0.5) dental_fees = [] # 遍历每个折叠面板区块 accordions = driver.find_elements(By.CLASS_NAME, 'accordion') for accordion in accordions: # 获取当前面板的标题 try: panel_title = accordion.find_element(By.CLASS_NAME, 'accordion-title').text.strip() except NoSuchElementException: panel_title = "未命名分类" # 获取当前面板下的所有项目行 content_block = accordion.find_element(By.CLASS_NAME, 'accordion-content') item_rows = content_block.find_elements(By.CSS_SELECTOR, '.row') for row in item_rows: try: # 提取项目名称和价格 item_name = row.find_element(By.CLASS_NAME, 'col-md-8').text.strip() item_price = row.find_element(By.CLASS_NAME, 'col-md-4').text.strip() dental_fees.append({ "分类标题": panel_title, "项目名称": item_name, "价格": item_price }) except NoSuchElementException: # 跳过格式异常的行 continue driver.quit() # 打印结构化数据 for item in dental_fees: print(f"分类: {item['分类标题']} | 项目: {item['项目名称']} | 价格: {item['价格']}")
关键改进点说明
- 精准展开面板:通过
accordion-header定位每个面板的头部,再找到对应的展开按钮,只点击需要展开的面板(判断图标文本为expand_more) - 区块内元素定位:遍历每个
accordion区块时,使用accordion.find_element来定位当前区块下的内容,确保获取的是对应分类下的项目 - 结构化提取:拆分每行的项目名称和价格,整理成字典格式,方便后续处理和存储
- 异常处理:增加了Cookie按钮不存在、点击被遮挡、元素缺失等情况的异常处理,提升代码稳定性
内容的提问来源于stack exchange,提问作者Ali Mohamed
相关产品推荐
相关产品推荐

