You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用Cypress将指定HTML表格转换为目标JSON格式?

实现HTML表格转指定格式JSON的方法

下面提供两种适配不同场景的实用实现方案:

方法一:前端JavaScript(浏览器环境)

如果需要在浏览器中直接处理页面上的表格,可通过原生JS遍历DOM元素生成目标JSON:

function tableToJson(table) {
  // 提取表头列名
  const headers = Array.from(table.querySelectorAll('thead th')).map(th => th.textContent.trim());
  // 遍历tbody生成每行的数据对象
  const rows = Array.from(table.querySelectorAll('tbody tr')).map(tr => {
    const rowValues = Array.from(tr.querySelectorAll('td')).map(td => td.textContent.trim());
    // 将列名与对应单元格值配对为对象
    return headers.reduce((obj, header, idx) => {
      obj[header] = rowValues[idx];
      return obj;
    }, {});
  });
  // 返回指定结构的JSON对象
  return { myrows: rows };
}

// 使用示例:获取页面第一个表格并转换
const targetTable = document.querySelector('table');
const resultJson = tableToJson(targetTable);
// 格式化输出结果
console.log(JSON.stringify(resultJson, null, 2));

逻辑说明:

  1. 先提取表头所有列的文本作为JSON对象的键名
  2. 遍历tbody内的每一行,提取单元格文本作为对应键的值
  3. 通过reduce将每行的键值对组装为对象,最终放入myrows数组中

方法二:Python脚本(后端/离线处理)

如果需要离线处理HTML文本,可使用Python结合BeautifulSoup库解析转换:

首先安装依赖:

pip install beautifulsoup4

转换代码:

from bs4 import BeautifulSoup
import json

def html_table_to_json(html_content):
    soup = BeautifulSoup(html_content, 'html.parser')
    table = soup.find('table')
    
    # 提取表头列名
    headers = [th.get_text(strip=True) for th in table.find('thead').find_all('th')]
    rows = []
    
    # 遍历处理每一行数据
    for tr in table.find('tbody').find_all('tr'):
        row_values = [td.get_text(strip=True) for td in tr.find_all('td')]
        row_obj = dict(zip(headers, row_values))
        rows.append(row_obj)
    
    # 构造目标JSON结构并返回格式化字符串
    return json.dumps({"myrows": rows}, indent=2, ensure_ascii=False)

# 使用示例:传入HTML文本
html_text = """
<table>
  <thead>
    <tr>
      <th>Column 1</th>
      <th>Column 2</th>
      <th>Column 3</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td>A1</td>
      <td>A2</td>
      <td>A3</td>
    </tr>
    <tr>
      <td>B1</td>
      <td>B2</td>
      <td>B3</td>
    </tr>
    <tr>
      <td>C1</td>
      <td>C2</td>
      <td>C3</td>
    </tr>
  </tbody>
</table>
"""
print(html_table_to_json(html_text))

逻辑说明:

  1. 用BeautifulSoup解析HTML文本,定位目标表格
  2. 提取表头文本和每行单元格文本
  3. 通过zip将表头与行数据配对为字典,最终组装成指定的JSON结构

通用注意事项

  • 两种方案都支持任意行数的表格,只要表头与每行的单元格数量对应即可
  • 若表格包含嵌套标签、图片等复杂内容,可根据需求修改文本提取逻辑(比如获取元素属性值、特定子标签内容)

内容的提问来源于stack exchange,提问作者Vivek Singh

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.07 19:55:30