如何用Python保留Word文档样式(粗体/斜体)并提取脚注存入SQL Server?
问题描述
我正在编写Python脚本,从Word文档(.docx)提取内容并插入到SQL Server数据库中。当前需求为保留粗体、斜体等文本样式,同时处理换行与脚注。目前使用python-docx库处理文档,已通过<br>成功转换换行,但文本样式(粗体/斜体)和脚注未出现在输出中。
已尝试的操作:
- 文本样式处理:遍历
paragraph.runs检测run.bold和run.italic,但带样式的文本未出现在数据库输出中。 - 脚注提取:通过自定义函数使用
doc.footnotes或检查“Footnote Text”样式提取脚注,函数无报错但脚注未出现在最终输出中。
代码片段:
处理样式的代码
text_with_style = [] if paragraph.runs: for run in paragraph.runs: styled_text = run.text.strip() if run.bold: styled_text = f"<b>{styled_text}</b>" if run.italic: styled_text = f"<i>{styled_text}</i>" text_with_style.append(styled_text) formatted_text = " ".join(text_with_style).replace("\n", "<br>")
提取脚注的代码
def extract_footnotes(doc): footnotes_text = [] if hasattr(doc, 'footnotes'): for footnote in doc.footnotes: footnotes_text.append(footnote.text.strip()) return footnotes_text
请问我遗漏了什么?如何可靠保留粗体/斜体样式并提取脚注,使其能被插入到SQL Server中?恳请提供建议或可运行示例。
解决方案
一、修复文本样式处理问题
现有代码存在两个核心问题:
strip()导致样式内容丢失:直接对run.text执行strip()会移除文本前后空格,可能丢失带样式的空白内容,或导致拼接后文本逻辑断连。- 样式叠加逻辑错误:若一个run同时包含粗体+斜体,现有代码会先添加
<b>标签,再被<i>标签覆盖,最终仅保留斜体样式。
修正后的样式处理代码:
def process_paragraph_style(paragraph): text_with_style = [] for run in paragraph.runs: styled_text = run.text # 保留原始文本,避免strip()丢失内容 # 处理多样式叠加 if run.bold and run.italic: styled_text = f"<b><i>{styled_text}</i></b>" elif run.bold: styled_text = f"<b>{styled_text}</b>" elif run.italic: styled_text = f"<i>{styled_text}</i>" text_with_style.append(styled_text) # 替换换行符为<br>,兼容不同换行格式 formatted_text = "".join(text_with_style).replace("\r\n", "<br>").replace("\n", "<br>") return formatted_text
二、修复脚注提取问题
footnote.text仅提取纯文本,会丢失样式,且默认包含脚注标记(如数字编号)。需遍历脚注的段落和run来保留样式,同时过滤标记内容:
def extract_footnotes_with_style(doc): footnotes = [] if hasattr(doc, 'footnotes'): for idx, footnote in enumerate(doc.footnotes): footnote_content = [] # 跳过第一个段落(通常是脚注标记,如[1]) for para in footnote.paragraphs[1:]: para_text = process_paragraph_style(para) footnote_content.append(para_text) # 给脚注添加编号,方便对应原文 footnotes.append(f"[#{idx+1}] {'<br>'.join(footnote_content)}") return "<br><br>".join(footnotes)
三、插入SQL Server的注意事项
- 字段类型选择:存储带HTML样式的文本,建议使用
NVARCHAR(MAX)或VARCHAR(MAX),避免长度限制。 - 参数化查询防注入:插入前必须转义SQL特殊字符(如单引号
'),使用pyodbc的参数化查询可自动处理该问题:
import pyodbc # 建立SQL Server连接 conn = pyodbc.connect('DRIVER={ODBC Driver 17 for SQL Server};SERVER=你的服务器地址;DATABASE=你的数据库名;UID=用户名;PWD=密码') cursor = conn.cursor() # 参数化插入语句 insert_query = """ INSERT INTO doc_content (body_content, footnotes) VALUES (?, ?) """ # 传入处理后的正文和脚注 cursor.execute(insert_query, (processed_content, final_footnotes)) conn.commit() cursor.close() conn.close()
完整可运行示例
from docx import Document import pyodbc def process_paragraph_style(paragraph): text_with_style = [] for run in paragraph.runs: styled_text = run.text if run.bold and run.italic: styled_text = f"<b><i>{styled_text}</i></b>" elif run.bold: styled_text = f"<b>{styled_text}</b>" elif run.italic: styled_text = f"<i>{styled_text}</i>" text_with_style.append(styled_text) formatted_text = "".join(text_with_style).replace("\r\n", "<br>").replace("\n", "<br>") return formatted_text def extract_footnotes_with_style(doc): footnotes = [] if hasattr(doc, 'footnotes'): for idx, footnote in enumerate(doc.footnotes): footnote_content = [] for para in footnote.paragraphs[1:]: para_text = process_paragraph_style(para) footnote_content.append(para_text) footnotes.append(f"[#{idx+1}] {'<br>'.join(footnote_content)}") return "<br><br>".join(footnotes) def main(): # 加载目标Word文档 doc = Document("你的文档路径.docx") # 处理正文内容,跳过空段落 processed_content = [] for para in doc.paragraphs: para_text = process_paragraph_style(para) if para_text.strip(): processed_content.append(para_text) final_content = "<br><br>".join(processed_content) # 处理脚注 final_footnotes = extract_footnotes_with_style(doc) # 插入SQL Server conn = pyodbc.connect('DRIVER={ODBC Driver 17 for SQL Server};SERVER=你的服务器地址;DATABASE=你的数据库名;UID=用户名;PWD=密码') cursor = conn.cursor() cursor.execute("INSERT INTO doc_storage (body_content, footnotes) VALUES (?, ?)", (final_content, final_footnotes)) conn.commit() print("数据插入成功") cursor.close() conn.close() if __name__ == "__main__": main()
内容的提问来源于stack exchange,提问作者Zahra
相关产品推荐
相关产品推荐

