Python 2.7 Selenium无法读取Java站点表格内容求助
问题分析与解决方案
先揪出最直接的问题:你代码里的表格ID拼写错误了!
对比一下:
- 你代码里写的是:
RAD_SPLIITER_dnn_ctr956_View_RadSpliter1 - 实际页面源码里的ID是:
RAD_SPLITTER_dnn_ctr956_View_RadSplitter1
两处拼写错误:
- 第一个单词:
SPLIITER→ 正确是SPLITTER(多了一个I,少了一个T) - 最后一个单词:
RadSpliter→ 正确是RadSplitter(少了一个T)
这绝对是导致你找不到表格的首要原因!
除此之外,还有几个常见的坑需要排查:
1. 表格是否在iframe中?
很多基于Telerik控件(比如这个RadSplitter)的Java站点,会把内容放在iframe里。如果表格藏在iframe内,你必须先切换到对应的iframe才能定位元素。可以这样检查:
- 打开浏览器开发者工具(F12),找到表格元素,看它的父级是否包含
<iframe>标签。 - 如果存在iframe,在代码里先切换过去,示例:
# 可以通过ID、name或xpath定位iframe iframe = WebDriverWait(driver, 10).until( EC.presence_of_element_located((By.XPATH, "//iframe[contains(@src, 'Oddsmatcher')]")) ) driver.switch_to.frame(iframe) # 之后再查找表格
2. 等待策略不够可靠
你用了implicitly_wait和固定时长的time.sleep,但这类方式不够灵活——有时候元素已经存在但还没渲染完成,或者动态加载需要更长时间。推荐用显式等待来确保元素可被定位:
from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC # 等待表格出现,最长等待15秒 table = WebDriverWait(driver, 15).until( EC.presence_of_element_located((By.ID, "RAD_SPLITTER_dnn_ctr956_View_RadSplitter1")) )
3. 查找行元素的路径错误
你在查找表格行时用了table.find_element_by_xpath('tr'),这是绝对路径查找,会从整个文档根节点开始找,而不是在表格内部。应该改成相对路径.//tr,这样才会限定在当前表格范围内查找。
修正后的完整代码
把以上要点整合后,你的代码可以调整成这样:
from bs4 import BeautifulSoup from selenium import webdriver from selenium.webdriver.common.keys import Keys from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC import time import sys, os, requests from os import system def main(): file = open("wbSc2.txt","w") print 'starting...' print >> file, 'starting...' site2 = "https://www.oddsmonkey.com/Tools/Oddsmatcher.aspx" driver = webdriver.Firefox() print 'grabbing site' print >> file, 'grabbing site' driver.get(site2) # 显式等待用户名输入框加载完成 user = WebDriverWait(driver, 10).until( EC.presence_of_element_located((By.ID, "dnn_ctr433_Login_Login_DNN_txtUsername")) ) password = driver.find_element_by_id("dnn_ctr433_Login_Login_DNN_txtPassword") user.send_keys('myusername') password.send_keys('mypassword') submit = driver.find_element_by_id("dnn_ctr433_Login_Login_DNN_cmdLogin") submit.click() # 等待弹窗关闭按钮可点击并关闭 close = WebDriverWait(driver, 10).until( EC.element_to_be_clickable((By.XPATH, "//button[@class='rltbActionButton rltbCloseButton']")) ) close.click() # 可选:如果表格在iframe中,解开下面的注释并调整定位逻辑 # try: # iframe = WebDriverWait(driver, 10).until( # EC.presence_of_element_located((By.XPATH, "//iframe[contains(@id, 'RadSplitter')]")) # ) # driver.switch_to.frame(iframe) # print "Switched to iframe" # print >> file, "Switched to iframe" # except: # print "No iframe found, proceeding directly" # print >> file, "No iframe found, proceeding directly" try: print 'attempting to find the table' print >> file, 'attempting to find the table' # 修正ID并显式等待表格加载 table = WebDriverWait(driver, 15).until( EC.presence_of_element_located((By.ID, "RAD_SPLITTER_dnn_ctr956_View_RadSplitter1")) ) print 'successfully found table' print >> file, 'successfully found table' print table.text print >> file, table.text except Exception as e: print 'failed to find table:', str(e) print >> file, 'failed to find table:', str(e) try: print 'attempting to find row' print >> file, 'attempting to find row' # 用相对路径查找表格内的行 row = table.find_element_by_xpath('.//tr') print 'successfully found row' print >> file, 'successfully found row' print row.text print >> file, row.text except Exception as e: print 'failed to find row:', str(e) print >> file, 'failed to find row:', str(e) # 可选:如果切换过iframe,记得切回主文档 # driver.switch_to.default_content() driver.close() system("pause") main()
先优先修正ID的拼写错误,这应该能解决大部分问题。如果还是找不到表格,再检查是否存在iframe,或者调整显式等待的时长。
内容的提问来源于stack exchange,提问作者1QuickQuestion
相关产品推荐
相关产品推荐

