如何使用Selenium定位p标签元素并解决文本提取报错问题
代码修正方案
现有代码问题梳理
- 执行搜索操作后未等待结果加载完成,直接查找元素会触发元素不存在报错
- XPath使用错误:
//p[@id='phoneDiv_80863']中的id为单条数据的专属标识,无法匹配所有结果,且子元素查询时XPath开头加//会从全局文档检索,而非限定在当前遍历的row节点内 - 语法过时:Selenium 4+ 已废弃
find_element_by_xpath写法,需统一使用find_element(By.XPATH, 路径)格式 - 方法调用错误:Selenium获取元素文本的属性为
.text,.get_text()是BeautifulSoup库的方法,混用会触发属性不存在报错 - 路径转义问题:Windows本地路径直接写反斜杠会被Python识别为转义符,需加
r前缀转换为原始字符串,同时ChromeDriver初始化现在推荐使用Service参数规范写法
修正后完整代码
from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.common.by import By from selenium.webdriver.support import expected_conditions as EC from selenium.webdriver.support.select import Select from selenium import webdriver from selenium.webdriver.chrome.service import Service # 规范ChromeDriver初始化写法,加r前缀解决路径转义问题 service = Service(executable_path=r'C:\Program Files (x86)\chromedriver.exe') driver = webdriver.Chrome(service=service) driver.maximize_window() wait = WebDriverWait(driver, 30) driver.get("https://www.counselingcalifornia.com/Find-a-Therapist") # 切换到内容iframe逻辑保留 wait.until(EC.frame_to_be_available_and_switch_to_it((By.CSS_SELECTOR, "iframe[id$='IFrame_htmIFrame']"))) select = Select(wait.until(EC.visibility_of_element_located((By.ID, "language_field")))) select.select_by_value('ENG') # 点击搜索 wait.until(EC.element_to_be_clickable((By.CSS_SELECTOR, "a#searchBtn"))).click() # 新增等待:搜索结果加载完成后再执行提取逻辑 wait.until(EC.visibility_of_all_elements_located((By.XPATH, "//div[@class='row']"))) dunk = driver.find_elements(By.XPATH, "//div[@class='row']") for dun in dunk: try: # 用相对路径匹配当前条目下的所有电话标签,starts-with适配动态id phone = dun.find_element(By.XPATH, ".//p[starts-with(@id,'phoneDiv_')]").text print(phone.strip()) except: # 跳过无电话的异常条目,避免脚本中断 continue # 执行完后关闭驱动 driver.quit()
内容的提问来源于stack exchange,提问作者Amen Aziz
相关产品推荐
相关产品推荐

