如何用BeautifulSoup将指定HTML表格转换为目标JSON格式?
问题:将HTML表格转换为指定结构的JSON字典
原始HTML表格
<table> <tbody> <tr> <th class="left" colspan="7"> <p>Some text</p> </th> </tr> <tr> <td class="left print-wide" colspan="2"> </td> <td class="print-wide" colspan="13">some-text</td> </tr> <tr> <td class="left"><br /></td> <td><strong>ABC </strong></td> <td><strong>≤25%</strong></td> <td><strong>≤75%</strong></td> <td><strong>≤100%</strong></td> </tr> <tr> <td class="left">1 month</td> <td>3,93%</td> <td>4,05%</td> <td>4,09%</td> <td>4,18%</td> </tr> <tr> <td class="left">3 months</td> <td>4,12%</td> <td>4,24%</td> <td>4,28%</td> <td>4,37%</td> </tr> <tr> <td class="left">6 months</td> <td>4,23%</td> <td>4,35%</td> <td>4,39%</td> <td>4,48%</td> </tr> </tbody> </table>
目标JSON结构
{ "1 month": { "ABC": "3,93%", "≤25%": "4,05%", "≤75%": "4,09%", "≤100%": "4,18%" }, "3 month": { "ABC": "4,12%", "≤25%": "4,24%", "≤75%": "4,28%", "≤100%": "4,37%" }, "6 month": { "ABC": "4,23%", "≤25%": "4,35%", "≤75%": "4,39%", "≤100%": "4,48%" } }
已完成的代码
已成功提取月份列表:
from bs4 import BeautifulSoup soup = BeautifulSoup(body, "html.parser") table = soup.find("table") headers = [header.text.strip() for header in table.find_all('td', class_="left")] del headers[:2] print(headers) # 输出: ['1 month', '3 months', '6 months']
完整解决方案
要构建目标JSON,需先提取列标题,再遍历数据行匹配对应值:
from bs4 import BeautifulSoup import json # 假设body是你的HTML内容 soup = BeautifulSoup(body, "html.parser") table = soup.find("table") tbody = table.find("tbody") rows = tbody.find_all("tr") # 提取列标题(ABC、≤25%、≤75%、≤100%) header_row = rows[2] # 对应包含strong标签的行 column_headers = [cell.text.strip() for cell in header_row.find_all("td") if cell.text.strip()] # 提取数据行(第3、4、5行,索引从0开始) data_rows = rows[3:6] result = {} for idx, row in enumerate(data_rows): # 处理月份键,将"3 months"转为"3 month" month_key = headers[idx].replace("months", "month") if "months" in headers[idx] else headers[idx] # 获取当前行的所有数值 values = [cell.text.strip() for cell in row.find_all("td") if cell.text.strip()] # 构建子字典,列标题对应数值(跳过第一个值,即月份) result[month_key] = dict(zip(column_headers, values[1:])) # 输出格式化后的JSON print(json.dumps(result, indent=4, ensure_ascii=False))
代码说明:
- 定位到包含列标题的行,提取并清洗标题文本
- 取后三行作为数据行,遍历每行时:
- 调整月份文本格式,匹配目标JSON的键名
- 提取行内数值,用
zip将列标题与数值配对成子字典 - 将子字典添加到结果字典中,以处理后的月份为键
内容的提问来源于stack exchange,提问作者C-nan
相关产品推荐
相关产品推荐

