You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将HTML表格内容转换为Plotly Dash HTML组件字典?

问题:将Word转HTML表格转换为Dash组件字典

我用Mammoth工具将Word文件转换为HTML表格,表格包含近100个条目及其多段落定义。需要将其转换为Python字典:以条目名为键,值为html.Div包裹多个html.P组成的Dash HTML组件,方便按条目名搜索。

示例HTML表格片段

<tr>
        <td>
            <p>Item 1</p>
        </td>
        <td>
            <p>Definition for Item 1.</p>
            <p>This may contain several paragraphs.</p>
        </td>
</tr>
<tr>
        <td>
            <p>Item 2</p>
        </td>
        <td>
            <p>Definition for Item 2.</p>
            <p>This may contain several paragraphs.</p>
            <p>And another paragraph here.</p>
        </td>
</tr>

目标字典结构

items_dict = {
   'Item 1': html.Div([
        html.P("Definition for Item 1."),
        html.P("This may contain several paragraphs."),
    ]),
    'Item 2': html.Div([
        html.P("Definition for Item 2."),
        html.P("This may contain several paragraphs."),
        html.P("And another paragraph here."),
    ]),
}

解决方案

实现步骤

  • 使用BeautifulSoup解析HTML内容,遍历表格的每一行<tr>
  • 提取每行第一个<td>内的<p>文本作为字典的键(条目名)
  • 提取每行第二个<td>内的所有<p>文本,生成对应的html.P组件列表,再用html.Div包裹作为字典的值
  • 处理文本中的空白字符,保证内容整洁

完整代码示例

from bs4 import BeautifulSoup
from dash import html

# 替换为Mammoth转换得到的完整HTML表格内容
html_content = """
<table>
<tr>
        <td>
            <p>Item 1</p>
        </td>
        <td>
            <p>Definition for Item 1.</p>
            <p>This may contain several paragraphs.</p>
        </td>
</tr>
<tr>
        <td>
            <p>Item 2</p>
        </td>
        <td>
            <p>Definition for Item 2.</p>
            <p>This may contain several paragraphs.</p>
            <p>And another paragraph here.</p>
        </td>
</tr>
</table>
"""

# 解析HTML
soup = BeautifulSoup(html_content, "html.parser")
items_dict = {}

# 遍历所有表格行
for tr in soup.find_all("tr"):
    tds = tr.find_all("td")
    if len(tds) != 2:
        continue  # 跳过结构不符合的行
    
    # 提取条目名(第一个td下的p标签文本)
    item_name = tds[0].find("p").get_text(strip=True)
    
    # 提取所有定义段落,生成html.P组件列表
    definition_paragraphs = []
    for p in tds[1].find_all("p"):
        text = p.get_text(strip=True)
        if text:  # 跳过空段落
            definition_paragraphs.append(html.P(text))
    
    # 存入字典
    if item_name and definition_paragraphs:
        items_dict[item_name] = html.Div(definition_paragraphs)

# 验证结果
print(items_dict.keys())

注意事项

  • 如果HTML中存在嵌套格式标签(如<b>、<i>),可保留标签结构,将标签转换为对应的Dash组件传入html.P的子元素,而非仅提取纯文本
  • 若Mammoth转换的HTML包含额外冗余标签,可根据需求调整BeautifulSoup的提取逻辑
  • 针对近100条数据,该方案处理效率足够,无需额外优化

内容的提问来源于stack exchange,提问作者dpppstl

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.01 08:47:36