使用Selenium构建AI语音助手语音转文本时无法获取输出文本
问题解决:Selenium无法获取Speechnotes语音转文本输出
核心问题分析
你使用的绝对XPath/html/body/div[4]/div/div[1]/div[2]定位的并非实际存储语音转文本结果的容器,该元素仅为空白占位或非目标区域,导致获取到的文本仅为空白字符(长度3大概率是换行/空格)。
解决方案
Speechnotes的文本输出区域是一个带contenteditable="true"的div,且拥有唯一IDresults,用这个ID定位是最稳定的方式,无需依赖易变的绝对XPath。
修改后的完整代码
from time import sleep import warnings from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.chrome.service import Service from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC warnings.simplefilter("ignore") try: path = "c:\Users\Utente\Desktop\SPARK\chromedriver.exe" options = webdriver.ChromeOptions() options.headless = False options.add_experimental_option("excludeSwitches", ["enable-logging"]) options.add_argument("--log-level=3") options.add_argument("--use-fake-ui-for-media-stream") options.add_argument("--use-fake-device-for-media-stream") service = Service(path) driver = webdriver.Chrome(service=service, options=options) driver.get("https://speechnotes.co/dictate/") # 用显式等待替代sleep,提升稳定性 wait = WebDriverWait(driver, 10) # 选择意大利语(替换为相对XPath,避免绝对路径的脆弱性) it_option = wait.until(EC.element_to_be_clickable((By.XPATH, "//select[@id='languages']/option[@value='it-IT']"))) it_option.click() # 点击开始录音按钮 start_btn = wait.until(EC.element_to_be_clickable((By.XPATH, "//div[@class='start-btn']"))) start_btn.click() print("Listening...") except Exception as e: print(e) driver.quit() while True: try: # 定位正确的文本输出容器 text_element = wait.until(EC.presence_of_element_located((By.ID, "results"))) text = text_element.text.strip() stop = input("Type anything to stop --> ") print("识别结果:", text) if stop != "": driver.quit() break except Exception as e: print("获取文本出错:", e) break
关键优化点
- 替换元素定位方式:用
By.ID, "results"定位文本输出框,比绝对XPath更稳定,不受页面DOM结构变化影响。 - 改用显式等待:替代固定时长的
sleep,确保元素加载完成后再操作,避免因页面加载慢导致的元素未找到问题。 - 优化语言选择定位:用
value='it-IT'的option定位意大利语,比依赖option索引(42)更可靠,索引可能随网站更新变化。
内容的提问来源于stack exchange,提问作者fanti
相关产品推荐
相关产品推荐

