You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Python docx库时,正则匹配含问号的引号内容报错问题

问题根源

你代码报错的核心原因是:用re.findall提取引号内的内容后,直接把内容传给re.search做匹配——但引号内的内容如果包含正则元字符(比如?),re.search会把它当成正则语法解析,而非普通文本,导致匹配失败返回None,调用span()时就抛出AttributeError。

比如你举的例子“Hello? How are you?”,提取到的内容是Hello? How are you?,其中的?是正则里的非贪婪匹配标记,re.search会尝试用正则规则去匹配,自然找不到对应的文本,返回None。

修复方案

最可靠高效的方式是直接用re.finditer替代re.findall,它会返回所有匹配的Match对象,每个对象自带start()和end()方法,能直接拿到匹配内容的位置,完全避免二次搜索的问题:

from docx import Document
import re

document = Document(filepath)

def highlight_quotes(document):
    for paragraph in document.paragraphs:
        quotes = []
        # 遍历所有包含引号的匹配结果,每个match是Match对象
        for match in re.finditer(r'“(.*?)”', paragraph.text):
            # 如果你要的是包含前后引号的整个内容的位置:
            full_quote_start = match.start()
            full_quote_end = match.end()
            quotes.append((full_quote_start, full_quote_end))
            
            # 如果你只需要引号内文本的位置,就用分组1的索引:
            # inner_start = match.start(1)
            # inner_end = match.end(1)
            # quotes.append((inner_start, inner_end))
备选方案(不推荐)

如果一定要保留findall的写法,需要用re.escape()把提取到的内容转义,让正则把它当成普通文本处理,同时加个判断避免None的情况:

from docx import Document
import re

document = Document(filepath)

def highlight_quotes(document):
    for paragraph in document.paragraphs:
        if matches := re.findall(r'“(.*?)”', paragraph.text):
            quotes = []
            for content in matches:
                # 转义内容中的正则特殊字符
                location = re.search(re.escape(content), paragraph.text)
                if location:  # 防止匹配失败返回None
                    start_index, end_index = location.span()
                    quotes.append((start_index, end_index))

注意:这种方法有隐患——如果段落里有和引号内内容重复的文本,re.search会匹配到第一个出现的位置,而非原引号所在的位置,所以优先用finditer的方案。

内容的提问来源于stack exchange,提问作者feichangholls

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.17 10:17:22