基于Selenium与BeautifulSoup的网页爬取问题求助
Selenium+BeautifulSoup爬取表格问题排查与修复
1. 修复代码语法错误(缩进与try块未闭合)
你的代码存在缩进混乱、try块缺少except分支的语法问题,这会直接导致运行报错。修复后的完整代码如下:
from selenium import webdriver from selenium.webdriver.common.keys import Keys from bs4 import BeautifulSoup from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from selenium.webdriver.common.by import By # 初始化浏览器驱动(需提前配置对应浏览器驱动) driver = webdriver.Chrome() driver.get("目标页面的初始URL") # 补充目标页面入口地址 # 搜索操作 search_bar = driver.find_element(By.NAME, 'txtQuickSearch') search_param = input('Insert Search Parameter Here:') search_bar.send_keys(search_param) search_bar.send_keys(Keys.RETURN) # 显式等待表格加载 wait = WebDriverWait(driver, 30) wait.until(EC.presence_of_element_located((By.CLASS_NAME, 'testBlack'))) # 解析页面 try: soup = BeautifulSoup(driver.page_source, 'html.parser') table = soup.find('table', {'class': 'testBlack'}) if table: table_data = [] for row in table.find_all('tr'): row_data = [cell.text.strip() for cell in row.find_all('td')] table_data.append(row_data) print(table_data) else: print('Error: Could not find table') except Exception as e: print(f"运行报错: {str(e)}") finally: driver.quit()
2. 替换静态等待为显式等待
原代码用sleep(30)是固定时长等待,无法适配页面加载速度的波动。改用WebDriverWait显式等待,直到目标表格元素出现再继续执行,既保证可靠性又避免不必要的等待。
3. 排查表格是否在iframe中
页面源码包含x-frame-options="sameorigin",说明页面可能嵌套了iframe,表格可能在iframe内部。如果是这种情况,需要先切换到iframe再操作:
# 等待iframe加载完成并切换(需替换为实际iframe的定位属性,比如id/name) wait.until(EC.frame_to_be_available_and_switch_to_it((By.ID, "目标iframe的ID"))) # 之后再执行表格查找逻辑 # 操作完成后切回主文档 driver.switch_to.default_content()
4. 确认表格的定位属性是否正确
你提供的页面源码片段中未出现class="testBlack"的表格,可能是class名称错误,或表格是通过JavaScript动态渲染的。建议通过浏览器开发者工具(F12)查看实际渲染后的HTML,确认表格的正确class、id或XPath定位属性。
5. 升级元素定位API
原代码使用的find_element_by_name是旧版Selenium API,新版已推荐使用By类统一定位方式,避免API失效问题。
内容的提问来源于stack exchange,提问作者Victoria
相关产品推荐
相关产品推荐

