You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python结合docx将文本中字典值替换为对应键?

文档缩写替换问题:首次保留全称(缩写),后续替换为缩写

需求说明

需要使用python-docx处理Word文档,将字典中值(全称)替换为键(缩写),规则如下:

  • 全称首次出现时,保留「全称(缩写)」格式(保留原文本的大小写)
  • 后续所有出现的该全称,直接替换为缩写

示例:

字典:{'km': 'Kilometer'}
原文本:'The Kilometer (km) % is very powerful. Blah blah blah blah. The Kilometer this. The Kilometer (km) that km.', 'the Kilometer is called km'
转换后:'The Kilometer (km) % is very powerful. Blah blah blah blah. The km this. The km that km.', 'the km is called km'

原代码问题

第一段代码(直接全替换)

单纯全局替换所有全称为缩写,没有区分首次出现和后续出现的场景,完全不符合需求。

for paragraph in doc.paragraphs:
    for key, value in acronyms.items():
        # Replace the value with the key in the paragraph text
        paragraph.text = paragraph.text.replace(value, key)

第二段代码(尝试处理首次出现)

存在多处逻辑错误和代码问题:

  • 变量名拼写错误(qey、valves),且acronyms[keys]中keys未定义,直接运行会报错
  • 错误合并整个文档文本为单一字符串,破坏了原有的段落结构
  • 替换逻辑只覆盖了空格、句号的场景,未处理%、逗号等其他边界情况
  • 正则匹配的边界处理不严谨,容易出现误匹配或漏匹配
from docx import Document
import re

# Define your acronym dictionary
acronyms = {
    'km': 'kilometer',
    # Add more acronyms as needed
}

print(acronyms[keys])  # 错误:keys未定义
# Open the Word document
doc = Document("acronymnsTest.docx")

# Extract text from the document
document_text = ""

for paragraph in doc.paragraphs:
    document_text += paragraph.text + " "

# Split the document text into words
document_words = document_text.split()
document_words = ' '.join(document_words)

# Create a dictionary for first_used
first_used = {key: False for key in acronyms}

list1 = [document_words]
list2 = []

for qey, valves in acronyms.items():  # 变量名错误:应为key, value
    #print(valves)
    for inputText in list1:
        for key in first_used:
            if first_used[key] is False:
                replacement = "$%^#$$"
                if  valves + '(' + key + ')' in inputText:
                    inputText = re.sub(r'\b' + re.escape(valves) + r'(?=\s|%|\b)', replacement, inputText, count=1)
                    inputText = inputText.replace(' ' + valves + ' ', ' ' + key + ' ')
                    inputText = inputText.replace(' ' + valves + '.', ' ' + key + '.')
                    first_used[key] = True
                    print(inputText)
                    inputText = inputText.replace(replacement, valves)
                    list2.append(inputText)
                elif valves + ' ' + key + ' ' in inputText:
                    replacement = "$%^#$$"
                    inputText = re.sub(r'\b' + re.escape(valves + ' ' + key) + r'(?=\s|%|\b)', replacement, inputText, count=1)
                    inputText = inputText.replace(' ' + valves + ' ', ' ' + key + ' ')
                    inputText = inputText.replace(' ' + valves + '.', ' ' + key + '.')
                    inputText = inputText.replace(replacement, valves + ' (' + key + ')')
                    first_used[key] = True
                    list2.append(inputText)
                elif ' ' + key + ' ' in inputText:
                    replacement = "$%^#$$"
                    inputText = re.sub(r'\b' + re.escape(key) + r'\b', replacement, inputText, count=1)
                    inputText = inputText.replace(' ' + valves + ' ', ' ' + key + ' ')
                    inputText = inputText.replace(' ' + valves + '.', ' ' + key + '.')
                    inputText = inputText.replace(replacement, valves + ' (' + key + ')')
                    first_used[key] = True
                    list2.append(inputText)
                else:
                    pass
                    print("nope")
            else:
                inputText = inputText.replace(' ' + acronyms[key] + ' ', ' ' + key + ' ')
                list2.append(inputText)

# Print or use the modified text
for modified_text in list2:
    print(modified_text)

解决方案

以下代码可正确实现需求,严格遵循替换规则,同时处理各种边界情况:

from docx import Document
import re

# 定义缩写字典:键=缩写,值=全称(小写,用于不区分大小写匹配)
acronyms = {
    'km': 'kilometer',
    # 可添加更多条目,例如 'CPU': 'central processing unit'
}

# 标记每个缩写是否已完成首次出现的格式化
first_used = {key: False for key in acronyms}

# 打开目标文档
doc = Document("acronymnsTest.docx")

# 遍历每个段落处理文本
for paragraph in doc.paragraphs:
    current_text = paragraph.text
    for abbr, full_name in acronyms.items():
        if not first_used[abbr]:
            # 匹配第一个出现的全称(不区分大小写,保留原文本大小写)
            # 正则规则:匹配独立单词,后续接非单词字符或文本结尾
            pattern = re.compile(rf'\b(?i)({re.escape(full_name)})\b(?=\W|$)')
            match_result = pattern.search(current_text)
            if match_result:
                # 替换为「原全称(缩写)」格式,保留原大小写
                replacement_str = f"{match_result.group(0)} ({abbr})"
                current_text = pattern.sub(replacement_str, current_text, count=1)
                # 标记该缩写已完成首次格式化
                first_used[abbr] = True
        
        # 替换当前段落中剩余的所有全称为缩写(不区分大小写)
        replace_pattern = re.compile(rf'\b(?i){re.escape(full_name)}\b(?=\W|$)')
        current_text = replace_pattern.sub(abbr, current_text)
    
    # 更新段落文本
    paragraph.text = current_text

# 保存修改后的文档
doc.save("modified_acronyms.docx")

代码说明

  1. 正则匹配逻辑:
    • (?i):启用不区分大小写匹配,覆盖Kilometer、kilometer等各种大小写形式
    • \b:确保匹配独立单词,避免误匹配包含全称的其他词汇(例如不会匹配kilometer中的meter)
    • (?=\W|$):匹配全称后接非单词字符(标点、空格等)或文本结尾的场景,覆盖所有边界情况
  2. 首次出现处理:
    • 找到第一个匹配的全称后,保留原文本的大小写,拼接成「全称(缩写)」格式
    • 标记该缩写为已使用,后续段落直接替换所有全称为缩写
  3. 段落处理:
    • 逐个处理文档段落,保留原有的段落结构
    • 处理完成后直接更新段落文本,最后保存文档

内容的提问来源于stack exchange,提问作者ConfusedCarolinian

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.11 00:24:51