You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用Python解析网页表格?已获取table_rows但无法提取数据

提取网页表格数据的解决方案

嘿,看起来你已经搞定了最棘手的定位环节——找到目标表格并拿到所有行数据了!接下来提取具体数值其实很简单,我给你整理了完整的代码和步骤:

完整代码示例

options = webdriver.ChromeOptions()
options.add_argument('headless')
driver = webdriver.Chrome(r'C:\Users\wendle\BrowserDrivers\chromedriver.exe', options=options)
driver.get(wsr)
html = driver.page_source
soup = BeautifulSoup(html, 'lxml')
table = soup.find_all('table')
table_rows = table[2].find_all('tr')

# 初始化列表存储提取后的结构化数据
tool_records = []

# 遍历数据行:跳过前两行(第一行是DIFFUSION标题栏,第二行是表头)
for row in table_rows[2:]:
    # 获取当前行的所有单元格
    cells = row.find_all('td')
    # 逐个提取字段文本,strip()清理多余空格和换行
    tool_id = cells[0].get_text(strip=True)
    tool_type = cells[1].get_text(strip=True)
    status = cells[2].get_text(strip=True)
    date_time = cells[3].get_text(strip=True)
    minutes = cells[4].get_text(strip=True)
    employee = cells[5].get_text(strip=True)
    comments = cells[6].get_text(strip=True)
    
    # 将数据存入字典,结构清晰方便后续处理
    record = {
        'ToolId': tool_id,
        'Type': tool_type,
        'Status': status,
        'Date/Time': date_time,
        'Min': minutes,
        'Employee': employee,
        'Comments': comments
    }
    tool_records.append(record)

# 打印验证提取结果
for entry in tool_records:
    print(entry)

关键步骤说明

  • 跳过非数据行:table_rows[2:]是因为你的表格前两行是标题栏和表头,从第三行开始才是实际的工具数据,直接跳过这两行就能精准抓取目标内容。
  • 提取单元格文本:用find_all('td')拿到当前行的所有单元格,再通过索引cells[0]到cells[6]对应每个字段,get_text(strip=True)会自动去掉文本前后的空格、换行符,让数据更整洁。
  • 结构化存储:把每一行数据存入字典再添加到列表,后续不管是导出到Excel、CSV还是做数据分析,都能轻松处理。

如果之后需要调整字段或者处理更复杂的表格结构,随时调整就行!

内容的提问来源于stack exchange,提问作者james wendle

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.14 07:25:53