You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用Python访问并修改Word文档中的文本框内容?

如何用python-docx访问并修改Word文本框内容

在Word文档中,文本框不属于普通段落层级,而是以**形状(Shape)**的形式存储,python-docx的高层API没有直接提供访问接口,需要通过底层lxml操作XML节点来处理。

核心思路

Word中的文本框对应XML里的wps:txbxContent节点(命名空间为http://schemas.microsoft.com/office/word/2010/wordprocessingShape),我们需要遍历文档中所有的drawing元素,定位到这个节点后,就能像处理普通段落一样修改其中的文本内容。

完整代码示例

以下代码在你原有逻辑的基础上,新增了文本框的处理逻辑,将所有文本(包括文本框内的)替换为等长的'a'字符串:

from docx import Document
from docx.oxml.ns import qn

def replace_text_with_a(text):
    # 将文本替换为等长的'a'组成的字符串
    lst_of_words = text.split(' ')
    n_w_l = []
    for word in lst_of_words:
        n_w_l.append('a' * len(word))
    return ' '.join(n_w_l)

def process_paragraph_runs(paragraph):
    # 处理普通段落的run和图片,复用原有逻辑
    for run in paragraph.runs:
        if run.text:
            new_text = replace_text_with_a(run.text)
            run.text = new_text
            print(new_text)
        elif run._r.getchildren() and run._r.getchildren()[0].tag.endswith('drawing'):
            # 处理图片逻辑(原代码保留)
            drawing = run._r.getchildren()[0]
            inline = drawing.getchildren()[0]
            blip = inline.getchildren()[0]
            img_bytes = blip._blob
            width = inline.extent.cx
            height = inline.extent.cy
            new_run = paragraph.add_run()
            new_run.add_picture(img_bytes, width=width, height=height)
            paragraph._p.remove(run._r)

def process_textboxes(doc):
    # 遍历文档中所有的drawing元素,处理文本框
    namespace = 'http://schemas.microsoft.com/office/word/2010/wordprocessingShape'
    # 遍历正文部分的所有drawing
    for element in doc.element.body.iter():
        if element.tag == qn('w:drawing'):
            # 查找文本框内容节点
            txbx_content = element.find(qn(f'wps:txbxContent'))
            if txbx_content is not None:
                # 遍历文本框内的段落
                for p in txbx_content.findall(qn('w:p')):
                    # 构造docx的Paragraph对象
                    paragraph = doc._body._body._element_to_object(p)
                    process_paragraph_runs(paragraph)
    # 处理页眉页脚中的文本框(如果需要)
    for section in doc.sections:
        for header in section.headers:
            for element in header.element.iter():
                if element.tag == qn('w:drawing'):
                    txbx_content = element.find(qn(f'wps:txbxContent'))
                    if txbx_content is not None:
                        for p in txbx_content.findall(qn('w:p')):
                            paragraph = header._element_to_object(p)
                            process_paragraph_runs(paragraph)
        for footer in section.footers:
            for element in footer.element.iter():
                if element.tag == qn('w:drawing'):
                    txbx_content = element.find(qn(f'wps:txbxContent'))
                    if txbx_content is not None:
                        for p in txbx_content.findall(qn('w:p')):
                            paragraph = footer._element_to_object(p)
                            process_paragraph_runs(paragraph)

# 主流程
doc = Document('your_document.docx')

# 处理普通段落
for paragraph in doc.paragraphs:
    process_paragraph_runs(paragraph)

# 处理文本框
process_textboxes(doc)

# 保存修改后的文档
doc.save('modified_document.docx')

关键说明

  • qn()函数用于处理XML命名空间,避免直接写冗长的命名空间字符串
  • 文本框可能存在于正文、页眉、页脚等区域,代码中分别进行了遍历
  • replace_text_with_a()函数封装了文本替换逻辑,提高代码复用性

内容的提问来源于stack exchange,提问作者Zaks Hugh

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.29 17:42:06