You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用BeautifulSoup将指定HTML表格转换为目标JSON格式?

问题:将HTML表格转换为指定结构的JSON字典

原始HTML表格

<table>
  <tbody>
    <tr>
      <th class="left" colspan="7">
        <p>Some text</p>
      </th>
    </tr>
    <tr>
      <td class="left print-wide" colspan="2">  </td>
      <td class="print-wide" colspan="13">some-text</td>
    </tr>
    <tr>
      <td class="left"><br /></td>
      <td><strong>ABC   </strong></td>
      <td><strong>≤25%</strong></td>
      <td><strong>≤75%</strong></td>
      <td><strong>≤100%</strong></td>
    </tr>
    <tr>
      <td class="left">1 month</td>
      <td>3,93%</td>
      <td>4,05%</td>
      <td>4,09%</td>
      <td>4,18%</td>
    </tr>
    <tr>
      <td class="left">3 months</td>
      <td>4,12%</td>
      <td>4,24%</td>
      <td>4,28%</td>
      <td>4,37%</td>
    </tr>
    <tr>
      <td class="left">6 months</td>
      <td>4,23%</td>
      <td>4,35%</td>
      <td>4,39%</td>
      <td>4,48%</td>
    </tr>
  </tbody>
</table>

目标JSON结构

{
    "1 month": {
        "ABC": "3,93%",
        "≤25%": "4,05%",
        "≤75%": "4,09%",
        "≤100%": "4,18%"
    },
    "3 month": {
        "ABC": "4,12%",
        "≤25%": "4,24%",
        "≤75%": "4,28%",
        "≤100%": "4,37%"
    },
    "6 month": {
        "ABC": "4,23%",
        "≤25%": "4,35%",
        "≤75%": "4,39%",
        "≤100%": "4,48%"
    }
}

已完成的代码

已成功提取月份列表:

from bs4 import BeautifulSoup

soup = BeautifulSoup(body, "html.parser")
table = soup.find("table")
headers = [header.text.strip() for header in table.find_all('td', class_="left")]
del headers[:2]
print(headers)
# 输出: ['1 month', '3 months', '6 months']

完整解决方案

要构建目标JSON,需先提取列标题,再遍历数据行匹配对应值:

from bs4 import BeautifulSoup
import json

# 假设body是你的HTML内容
soup = BeautifulSoup(body, "html.parser")
table = soup.find("table")
tbody = table.find("tbody")
rows = tbody.find_all("tr")

# 提取列标题(ABC、≤25%、≤75%、≤100%)
header_row = rows[2]  # 对应包含strong标签的行
column_headers = [cell.text.strip() for cell in header_row.find_all("td") if cell.text.strip()]

# 提取数据行(第3、4、5行,索引从0开始)
data_rows = rows[3:6]

result = {}
for idx, row in enumerate(data_rows):
    # 处理月份键,将"3 months"转为"3 month"
    month_key = headers[idx].replace("months", "month") if "months" in headers[idx] else headers[idx]
    # 获取当前行的所有数值
    values = [cell.text.strip() for cell in row.find_all("td") if cell.text.strip()]
    # 构建子字典,列标题对应数值(跳过第一个值,即月份)
    result[month_key] = dict(zip(column_headers, values[1:]))

# 输出格式化后的JSON
print(json.dumps(result, indent=4, ensure_ascii=False))

代码说明:

  1. 定位到包含列标题的行,提取并清洗标题文本
  2. 取后三行作为数据行,遍历每行时:
    • 调整月份文本格式,匹配目标JSON的键名
    • 提取行内数值,用zip将列标题与数值配对成子字典
    • 将子字典添加到结果字典中,以处理后的月份为键

内容的提问来源于stack exchange,提问作者C-nan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.12 09:41:20