You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Selenium Python提取动态表格数据?代码无输出及存数求解

问题解决与数据存储实现

代码问题排查

  • URL错误:原代码使用的http://saasta.byu.edu/noauth/classSchedule/index.php并非目标站点,需替换为正确的https://commtech.byu.edu/noauth/classSchedule/index.php。
  • 重复搜索触发:同时用send_keys(Keys.RETURN)和点击搜索按钮会导致逻辑冲突,保留其中一种方式即可。
  • 表格行定位错误://tr会匹配页面所有<tr>元素,改为.//tr仅在目标表格内查找行,避免获取无关数据。
  • 等待条件优化:将presence_of_element_located替换为visibility_of_element_located,确保表格加载完成且可见,避免读取空内容。

修正后代码(含数据存储功能)

1. 依赖安装

先确保安装所需库:

pip install selenium

2. 完整代码

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.common.keys import Keys
from selenium.webdriver.chrome.service import Service
from selenium.webdriver.chrome.options import Options
import sqlite3

# 配置Chrome选项
c_options = Options()
c_options.add_experimental_option("detach", True)

# 配置ChromeDriver路径
s = Service('C:/Users/sidat/OneDrive/Desktop/python/WebDriver/chromedriver.exe')

# 目标站点URL
URL = "https://commtech.byu.edu/noauth/classSchedule/index.php"
driver = webdriver.Chrome(service=s, options=c_options)
driver.get(URL)

# 执行搜索操作
search_bar = driver.find_element(By.ID, "searchBar")
search_bar.send_keys("C S 142")
search_button = driver.find_element(By.ID, "searchBtn")
search_button.click()

# 等待表格加载完成并可见
table = WebDriverWait(driver, 15).until(
    EC.visibility_of_element_located((By.ID, "sectionTable"))
)

# 提取表格数据
rows = table.find_elements(By.XPATH, ".//tr")
class_data = []

for row in rows:
    cells = row.find_elements(By.TAG_NAME, "td")
    if cells:
        row_data = [cell.text.strip() for cell in cells]
        class_data.append(row_data)
        print("\t".join(row_data))

# ---------------------- 数据存储:写入文本文件 ----------------------
with open("cs142_class_schedule.txt", "w", encoding="utf-8") as f:
    # 写入表头
    headers = [th.text.strip() for th in table.find_elements(By.TAG_NAME, "th")]
    f.write("\t".join(headers) + "\n")
    # 写入内容行
    for row in class_data:
        f.write("\t".join(row) + "\n")
print("数据已成功写入cs142_class_schedule.txt")

# ---------------------- 数据存储:插入SQLite数据库 ----------------------
# 连接数据库(不存在则自动创建)
conn = sqlite3.connect("class_schedule.db")
cursor = conn.cursor()

# 创建表(可根据实际表头调整字段名)
create_table_sql = """
CREATE TABLE IF NOT EXISTS cs142_sections (
    crn TEXT,
    course_code TEXT,
    section TEXT,
    title TEXT,
    credits TEXT,
    instructor TEXT,
    meeting_time TEXT,
    location TEXT,
    status TEXT
)
"""
cursor.execute(create_table_sql)

# 批量插入数据
insert_sql = """
INSERT INTO cs142_sections VALUES (?, ?, ?, ?, ?, ?, ?, ?, ?)
"""
cursor.executemany(insert_sql, class_data)

# 提交并关闭连接
conn.commit()
conn.close()
print("数据已成功插入SQLite数据库class_schedule.db")

# 关闭浏览器
driver.quit()

关键说明

  • 文本存储:用制表符分隔数据,方便后续查看或导入其他工具。
  • SQL存储:示例用SQLite,若需用MySQL/PostgreSQL,只需修改数据库连接和插入逻辑。
  • 等待时长:设为15秒,适配网络较慢的情况,确保表格完全加载。

内容的提问来源于stack exchange,提问作者Sidath

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.08 23:01:09