使用Python Selenium操作Tulu词典网站遇NoSuchElementException问题求助
Tulu词典爬虫代码问题修复
问题概述
尝试使用Selenium从Tulu词典网站获取信息,代码运行时抛出NoSuchElementException,无法定位'Tulu'单选按钮,浏览器直接关闭。
原代码
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC # Create a new instance of the Chrome driver driver = webdriver.Chrome() # Open the webpage driver.get("https://tuludictionary.in/dictionary/cgi-bin/web/frame.html") # Select the 'Tulu' radio button tulu_radio = driver.find_element(By.XPATH, "/html/body/table/tbody/tr/td[1]/form/input[1]") tulu_radio.click() # Select the 'Anywhere' radio button anywhere_radio = driver.find_element(By.XPATH, "/html/body/table/tbody/tr/td[1]/form/input[6]") anywhere_radio.click() # Input "AA" in the text box search_box = driver.find_element(By.XPATH, "/html/body/table/tbody/tr/td[2]/input[1]") search_box.send_keys("AA") # Click on the search button search_button = driver.find_element(By.XPATH, "/html/body/table/tbody/tr/td[2]/input[2]") search_button.click() # Check and process search results output_area = driver.find_element(By.XPATH, "/html/body/table/tbody/tr/td") output_text = output_area.text.strip() # Wait for the results to load try: if output_text != "No results found": # Extract and print the text part of the links # Wait for the list of links to be displayed link_elements = WebDriverWait(driver, 10).until( EC.presence_of_all_elements_located((By.XPATH, "/html/body/table/tbody/tr/td")) ) if len(link_elements) > 0: # Get the text of the links links_text = [link.text for link in link_elements] print("List of links:") for link_text in links_text: print(link_text) # Click on the fourth link fourth_link = link_elements[3] fourth_link.click() else: print("NULL") result = None except: # Failed to retrieve results print("Failed to retrieve results") result = None # Close the browser driver.quit()
错误信息
File "C:\Users\chakr\AppData\Local\Temp\tempCodeRunnerFile.python", line 13, in <module> tulu_radio = driver.find_element(By.XPATH, "/html/body/table/tbody/tr/td[1]/form/input[1]") ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ File "C:\Program Files\Python311\Lib\site-packages\selenium\webdriver\remote\webdriver.py", line 740, in find_element return self.execute(Command.FIND_ELEMENT, {"using": by, "value": value})["value"] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ File "C:\Program Files\Python311\Lib\site-packages\selenium\webdriver\remote\webdriver.py", line 346, in execute self.error_handler.check_response(response) File "C:\Program Files\Python311\Lib\site-packages\selenium\webdriver\remote\errorhandler.py", line 245, in check_response raise exception_class(message, screen, stacktrace) selenium.common.exceptions.NoSuchElementException: Message: no such element: Unable to locate element: {"method":"xpath","selector":"/html/body/table/tbody/tr/td[1]/form/input[1]"} (Session info: chrome=114.0.5735.135); For documentation on this error, please visit: https://www.selenium.dev/documentation/webdriver/troubleshooting/errors#no-such-element-exception Stacktrace: Backtrace: GetHandleVerifier [0x00F5A813+48355] (No symbol) [0x00EEC4B1] (No symbol) [0x00DF5358] (No symbol) [0x00E209A5] (No symbol) [0x00E20B3B] (No symbol) [0x00E4E232] (No symbol) [0x00E3A784] (No symbol) [0x00E4C922] (No symbol) [0x00E3A536] (No symbol) [0x00E182DC] (No symbol) [0x00E193DD] GetHandleVerifier [0x011BAABD+2539405] GetHandleVerifier [0x011FA78F+2800735] GetHandleVerifier [0x011F456C+2775612] GetHandleVerifier [0x00FE51E0+616112] (No symbol) [0x00EF5F8C] (No symbol) [0x00EF2328] (No symbol) [0x00EF240B] (No symbol) [0x00EE4FF7] BaseThreadInitThunk [0x753C7D59+25] RtlInitializeExceptionChain [0x7725B74B+107] RtlClearBits [0x7725B6CF+191]
问题分析
- 页面使用iframe嵌套:目标页面的搜索表单实际在iframe内部,直接定位元素会找不到,必须先切换到iframe上下文。
- 绝对XPATH不稳定:使用从根节点开始的绝对路径定位元素,页面结构微小变化就会导致定位失败,应该改用相对XPATH或元素属性定位。
- 未等待元素加载:打开页面后直接定位元素,可能元素还未渲染完成,导致
NoSuchElementException。 - 结果处理逻辑错误:搜索结果区域的XPATH定位不准确,且等待时机错误,应该在点击搜索按钮后等待结果加载,而不是提前获取输出区域文本。
修复方案
修改后的代码
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC # 初始化Chrome驱动 driver = webdriver.Chrome() driver.maximize_window() # 打开目标页面 driver.get("https://tuludictionary.in/dictionary/cgi-bin/web/frame.html") try: # 切换到iframe(页面内容嵌套在iframe中) WebDriverWait(driver, 10).until( EC.frame_to_be_available_and_switch_to_it((By.NAME, "main")) ) # 等待并点击'Tulu'单选按钮(使用相对XPATH,按属性定位) tulu_radio = WebDriverWait(driver, 10).until( EC.element_to_be_clickable((By.XPATH, "//input[@value='tulu']")) ) tulu_radio.click() # 等待并点击'Anywhere'单选按钮 anywhere_radio = WebDriverWait(driver, 10).until( EC.element_to_be_clickable((By.XPATH, "//input[@value='any']")) ) anywhere_radio.click() # 等待并在搜索框输入"AA" search_box = WebDriverWait(driver, 10).until( EC.visibility_of_element_located((By.XPATH, "//input[@name='word']")) ) search_box.send_keys("AA") # 等待并点击搜索按钮 search_button = WebDriverWait(driver, 10).until( EC.element_to_be_clickable((By.XPATH, "//input[@value='Search']")) ) search_button.click() # 等待结果加载,检查是否有结果 output_area = WebDriverWait(driver, 10).until( EC.visibility_of_element_located((By.XPATH, "//td[@valign='top']")) ) output_text = output_area.text.strip() if output_text == "No results found": print("NULL") result = None else: # 获取所有结果链接 link_elements = WebDriverWait(driver, 10).until( EC.presence_of_all_elements_located((By.XPATH, "//td[@valign='top']//a")) ) print("List of links:") for idx, link in enumerate(link_elements, 1): print(f"{idx}. {link.text}") # 如果链接数量≥4,点击第四个链接 if len(link_elements) >= 4: fourth_link = link_elements[3] fourth_link.click() # 可选:等待新页面内容加载 WebDriverWait(driver, 10).until( EC.visibility_of_element_located((By.XPATH, "//table[@border='1']")) ) print("已点击第四个链接,页面加载完成") except Exception as e: print(f"出错了: {str(e)}") result = None finally: # 关闭浏览器 driver.quit()
关键修改点
- 切换iframe:使用
frame_to_be_available_and_switch_to_it等待并切换到名为main的iframe,这是定位内部元素的前提。 - 稳定定位方式:改用基于元素属性(如
value、name)的相对XPATH,避免绝对路径的脆弱性。 - 全局显式等待:所有元素操作前都用
WebDriverWait等待元素可点击/可见,确保元素渲染完成。 - 优化结果逻辑:先等待搜索结果加载,再判断是否有结果,准确定位结果链接的XPATH。
- 异常处理优化:捕获具体异常信息,便于调试,使用
finally确保浏览器关闭。
内容的提问来源于stack exchange,提问作者Milind
相关产品推荐
相关产品推荐

