使用Python docx库时,正则匹配含问号的引号内容报错问题
问题根源
你代码报错的核心原因是:用re.findall提取引号内的内容后,直接把内容传给re.search做匹配——但引号内的内容如果包含正则元字符(比如?),re.search会把它当成正则语法解析,而非普通文本,导致匹配失败返回None,调用span()时就抛出AttributeError。
比如你举的例子“Hello? How are you?”,提取到的内容是Hello? How are you?,其中的?是正则里的非贪婪匹配标记,re.search会尝试用正则规则去匹配,自然找不到对应的文本,返回None。
修复方案
最可靠高效的方式是直接用re.finditer替代re.findall,它会返回所有匹配的Match对象,每个对象自带start()和end()方法,能直接拿到匹配内容的位置,完全避免二次搜索的问题:
from docx import Document import re document = Document(filepath) def highlight_quotes(document): for paragraph in document.paragraphs: quotes = [] # 遍历所有包含引号的匹配结果,每个match是Match对象 for match in re.finditer(r'“(.*?)”', paragraph.text): # 如果你要的是包含前后引号的整个内容的位置: full_quote_start = match.start() full_quote_end = match.end() quotes.append((full_quote_start, full_quote_end)) # 如果你只需要引号内文本的位置,就用分组1的索引: # inner_start = match.start(1) # inner_end = match.end(1) # quotes.append((inner_start, inner_end))
备选方案(不推荐)
如果一定要保留findall的写法,需要用re.escape()把提取到的内容转义,让正则把它当成普通文本处理,同时加个判断避免None的情况:
from docx import Document import re document = Document(filepath) def highlight_quotes(document): for paragraph in document.paragraphs: if matches := re.findall(r'“(.*?)”', paragraph.text): quotes = [] for content in matches: # 转义内容中的正则特殊字符 location = re.search(re.escape(content), paragraph.text) if location: # 防止匹配失败返回None start_index, end_index = location.span() quotes.append((start_index, end_index))
注意:这种方法有隐患——如果段落里有和引号内内容重复的文本,re.search会匹配到第一个出现的位置,而非原引号所在的位置,所以优先用finditer的方案。
内容的提问来源于stack exchange,提问作者feichangholls
相关产品推荐
相关产品推荐

