如何用Python结合docx将文本中字典值替换为对应键?
文档缩写替换问题:首次保留全称(缩写),后续替换为缩写
需求说明
需要使用python-docx处理Word文档,将字典中值(全称)替换为键(缩写),规则如下:
- 全称首次出现时,保留「全称(缩写)」格式(保留原文本的大小写)
- 后续所有出现的该全称,直接替换为缩写
示例:
字典:
{'km': 'Kilometer'}
原文本:'The Kilometer (km) % is very powerful. Blah blah blah blah. The Kilometer this. The Kilometer (km) that km.', 'the Kilometer is called km'
转换后:'The Kilometer (km) % is very powerful. Blah blah blah blah. The km this. The km that km.', 'the km is called km'
原代码问题
第一段代码(直接全替换)
单纯全局替换所有全称为缩写,没有区分首次出现和后续出现的场景,完全不符合需求。
for paragraph in doc.paragraphs: for key, value in acronyms.items(): # Replace the value with the key in the paragraph text paragraph.text = paragraph.text.replace(value, key)
第二段代码(尝试处理首次出现)
存在多处逻辑错误和代码问题:
- 变量名拼写错误(
qey、valves),且acronyms[keys]中keys未定义,直接运行会报错 - 错误合并整个文档文本为单一字符串,破坏了原有的段落结构
- 替换逻辑只覆盖了空格、句号的场景,未处理%、逗号等其他边界情况
- 正则匹配的边界处理不严谨,容易出现误匹配或漏匹配
from docx import Document import re # Define your acronym dictionary acronyms = { 'km': 'kilometer', # Add more acronyms as needed } print(acronyms[keys]) # 错误:keys未定义 # Open the Word document doc = Document("acronymnsTest.docx") # Extract text from the document document_text = "" for paragraph in doc.paragraphs: document_text += paragraph.text + " " # Split the document text into words document_words = document_text.split() document_words = ' '.join(document_words) # Create a dictionary for first_used first_used = {key: False for key in acronyms} list1 = [document_words] list2 = [] for qey, valves in acronyms.items(): # 变量名错误:应为key, value #print(valves) for inputText in list1: for key in first_used: if first_used[key] is False: replacement = "$%^#$$" if valves + '(' + key + ')' in inputText: inputText = re.sub(r'\b' + re.escape(valves) + r'(?=\s|%|\b)', replacement, inputText, count=1) inputText = inputText.replace(' ' + valves + ' ', ' ' + key + ' ') inputText = inputText.replace(' ' + valves + '.', ' ' + key + '.') first_used[key] = True print(inputText) inputText = inputText.replace(replacement, valves) list2.append(inputText) elif valves + ' ' + key + ' ' in inputText: replacement = "$%^#$$" inputText = re.sub(r'\b' + re.escape(valves + ' ' + key) + r'(?=\s|%|\b)', replacement, inputText, count=1) inputText = inputText.replace(' ' + valves + ' ', ' ' + key + ' ') inputText = inputText.replace(' ' + valves + '.', ' ' + key + '.') inputText = inputText.replace(replacement, valves + ' (' + key + ')') first_used[key] = True list2.append(inputText) elif ' ' + key + ' ' in inputText: replacement = "$%^#$$" inputText = re.sub(r'\b' + re.escape(key) + r'\b', replacement, inputText, count=1) inputText = inputText.replace(' ' + valves + ' ', ' ' + key + ' ') inputText = inputText.replace(' ' + valves + '.', ' ' + key + '.') inputText = inputText.replace(replacement, valves + ' (' + key + ')') first_used[key] = True list2.append(inputText) else: pass print("nope") else: inputText = inputText.replace(' ' + acronyms[key] + ' ', ' ' + key + ' ') list2.append(inputText) # Print or use the modified text for modified_text in list2: print(modified_text)
解决方案
以下代码可正确实现需求,严格遵循替换规则,同时处理各种边界情况:
from docx import Document import re # 定义缩写字典:键=缩写,值=全称(小写,用于不区分大小写匹配) acronyms = { 'km': 'kilometer', # 可添加更多条目,例如 'CPU': 'central processing unit' } # 标记每个缩写是否已完成首次出现的格式化 first_used = {key: False for key in acronyms} # 打开目标文档 doc = Document("acronymnsTest.docx") # 遍历每个段落处理文本 for paragraph in doc.paragraphs: current_text = paragraph.text for abbr, full_name in acronyms.items(): if not first_used[abbr]: # 匹配第一个出现的全称(不区分大小写,保留原文本大小写) # 正则规则:匹配独立单词,后续接非单词字符或文本结尾 pattern = re.compile(rf'\b(?i)({re.escape(full_name)})\b(?=\W|$)') match_result = pattern.search(current_text) if match_result: # 替换为「原全称(缩写)」格式,保留原大小写 replacement_str = f"{match_result.group(0)} ({abbr})" current_text = pattern.sub(replacement_str, current_text, count=1) # 标记该缩写已完成首次格式化 first_used[abbr] = True # 替换当前段落中剩余的所有全称为缩写(不区分大小写) replace_pattern = re.compile(rf'\b(?i){re.escape(full_name)}\b(?=\W|$)') current_text = replace_pattern.sub(abbr, current_text) # 更新段落文本 paragraph.text = current_text # 保存修改后的文档 doc.save("modified_acronyms.docx")
代码说明
- 正则匹配逻辑:
(?i):启用不区分大小写匹配,覆盖Kilometer、kilometer等各种大小写形式\b:确保匹配独立单词,避免误匹配包含全称的其他词汇(例如不会匹配kilometer中的meter)(?=\W|$):匹配全称后接非单词字符(标点、空格等)或文本结尾的场景,覆盖所有边界情况
- 首次出现处理:
- 找到第一个匹配的全称后,保留原文本的大小写,拼接成「全称(缩写)」格式
- 标记该缩写为已使用,后续段落直接替换所有全称为缩写
- 段落处理:
- 逐个处理文档段落,保留原有的段落结构
- 处理完成后直接更新段落文本,最后保存文档
内容的提问来源于stack exchange,提问作者ConfusedCarolinian
相关产品推荐
相关产品推荐

