You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python保留Word文档样式(粗体/斜体)并提取脚注存入SQL Server?

问题描述

我正在编写Python脚本,从Word文档(.docx)提取内容并插入到SQL Server数据库中。当前需求为保留粗体、斜体等文本样式,同时处理换行与脚注。目前使用python-docx库处理文档,已通过<br>成功转换换行,但文本样式(粗体/斜体)和脚注未出现在输出中。

已尝试的操作:

  • 文本样式处理:遍历paragraph.runs检测run.bold和run.italic,但带样式的文本未出现在数据库输出中。
  • 脚注提取:通过自定义函数使用doc.footnotes或检查“Footnote Text”样式提取脚注,函数无报错但脚注未出现在最终输出中。

代码片段:

处理样式的代码

text_with_style = []
if paragraph.runs:
    for run in paragraph.runs:
        styled_text = run.text.strip()
        if run.bold:
            styled_text = f"<b>{styled_text}</b>"
        if run.italic:
            styled_text = f"<i>{styled_text}</i>"
        text_with_style.append(styled_text)

formatted_text = " ".join(text_with_style).replace("\n", "<br>")

提取脚注的代码

def extract_footnotes(doc):
    footnotes_text = []
    if hasattr(doc, 'footnotes'):
        for footnote in doc.footnotes:
            footnotes_text.append(footnote.text.strip())
    return footnotes_text

请问我遗漏了什么?如何可靠保留粗体/斜体样式并提取脚注,使其能被插入到SQL Server中?恳请提供建议或可运行示例。


解决方案

一、修复文本样式处理问题

现有代码存在两个核心问题:

  1. strip()导致样式内容丢失:直接对run.text执行strip()会移除文本前后空格,可能丢失带样式的空白内容,或导致拼接后文本逻辑断连。
  2. 样式叠加逻辑错误:若一个run同时包含粗体+斜体,现有代码会先添加<b>标签,再被<i>标签覆盖,最终仅保留斜体样式。

修正后的样式处理代码:

def process_paragraph_style(paragraph):
    text_with_style = []
    for run in paragraph.runs:
        styled_text = run.text  # 保留原始文本,避免strip()丢失内容
        # 处理多样式叠加
        if run.bold and run.italic:
            styled_text = f"<b><i>{styled_text}</i></b>"
        elif run.bold:
            styled_text = f"<b>{styled_text}</b>"
        elif run.italic:
            styled_text = f"<i>{styled_text}</i>"
        text_with_style.append(styled_text)
    # 替换换行符为<br>,兼容不同换行格式
    formatted_text = "".join(text_with_style).replace("\r\n", "<br>").replace("\n", "<br>")
    return formatted_text

二、修复脚注提取问题

footnote.text仅提取纯文本,会丢失样式,且默认包含脚注标记(如数字编号)。需遍历脚注的段落和run来保留样式,同时过滤标记内容:

def extract_footnotes_with_style(doc):
    footnotes = []
    if hasattr(doc, 'footnotes'):
        for idx, footnote in enumerate(doc.footnotes):
            footnote_content = []
            # 跳过第一个段落(通常是脚注标记,如[1])
            for para in footnote.paragraphs[1:]:
                para_text = process_paragraph_style(para)
                footnote_content.append(para_text)
            # 给脚注添加编号,方便对应原文
            footnotes.append(f"[#{idx+1}] {'<br>'.join(footnote_content)}")
    return "<br><br>".join(footnotes)

三、插入SQL Server的注意事项

  1. 字段类型选择:存储带HTML样式的文本,建议使用NVARCHAR(MAX)或VARCHAR(MAX),避免长度限制。
  2. 参数化查询防注入:插入前必须转义SQL特殊字符(如单引号'),使用pyodbc的参数化查询可自动处理该问题:
import pyodbc

# 建立SQL Server连接
conn = pyodbc.connect('DRIVER={ODBC Driver 17 for SQL Server};SERVER=你的服务器地址;DATABASE=你的数据库名;UID=用户名;PWD=密码')
cursor = conn.cursor()

# 参数化插入语句
insert_query = """
INSERT INTO doc_content (body_content, footnotes)
VALUES (?, ?)
"""
# 传入处理后的正文和脚注
cursor.execute(insert_query, (processed_content, final_footnotes))
conn.commit()
cursor.close()
conn.close()

完整可运行示例

from docx import Document
import pyodbc

def process_paragraph_style(paragraph):
    text_with_style = []
    for run in paragraph.runs:
        styled_text = run.text
        if run.bold and run.italic:
            styled_text = f"<b><i>{styled_text}</i></b>"
        elif run.bold:
            styled_text = f"<b>{styled_text}</b>"
        elif run.italic:
            styled_text = f"<i>{styled_text}</i>"
        text_with_style.append(styled_text)
    formatted_text = "".join(text_with_style).replace("\r\n", "<br>").replace("\n", "<br>")
    return formatted_text

def extract_footnotes_with_style(doc):
    footnotes = []
    if hasattr(doc, 'footnotes'):
        for idx, footnote in enumerate(doc.footnotes):
            footnote_content = []
            for para in footnote.paragraphs[1:]:
                para_text = process_paragraph_style(para)
                footnote_content.append(para_text)
            footnotes.append(f"[#{idx+1}] {'<br>'.join(footnote_content)}")
    return "<br><br>".join(footnotes)

def main():
    # 加载目标Word文档
    doc = Document("你的文档路径.docx")
    
    # 处理正文内容,跳过空段落
    processed_content = []
    for para in doc.paragraphs:
        para_text = process_paragraph_style(para)
        if para_text.strip():
            processed_content.append(para_text)
    final_content = "<br><br>".join(processed_content)
    
    # 处理脚注
    final_footnotes = extract_footnotes_with_style(doc)
    
    # 插入SQL Server
    conn = pyodbc.connect('DRIVER={ODBC Driver 17 for SQL Server};SERVER=你的服务器地址;DATABASE=你的数据库名;UID=用户名;PWD=密码')
    cursor = conn.cursor()
    cursor.execute("INSERT INTO doc_storage (body_content, footnotes) VALUES (?, ?)", (final_content, final_footnotes))
    conn.commit()
    print("数据插入成功")
    cursor.close()
    conn.close()

if __name__ == "__main__":
    main()

内容的提问来源于stack exchange,提问作者Zahra

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.13 13:05:21