正则实现按标点拆分字符串时,如何避免误拆分缩写?
按标点拆分句子时正确保留缩写并避免错误拆分
需求
需要实现一个函数,按句号、感叹号、问号、分号、冒号拆分包含多句子的字符串,同时满足:
- 不拆分缩写(如
Univ.、Dept.、Prof.、et al.等) - 保留缩写中的点号,不能像之前参考的方案那样移除这类点号(比如把
U.S.A.转为USA)
现有实现代码
import re def split_string_by_punctuation(line: str) -> list[str]: """ Splits a given string into a list of strings using terminal punctuation marks (., !, ?, or :) as delimiters. This function utilizes regular expression patterns to ensure that abbreviations, honorifics, and certain special cases are not considered as sentence delimiters. Args: line (str): The input string to be split into sentences. Returns: list: A list of strings representing the sentences obtained after splitting the input string. Notes: - Negative lookbehind is used to exclude abbreviations (e.g., "e.g.", "i.e.", "U.S.A."), which might have a period but are not the end of a sentence. - Negative lookbehind is also used to exclude honorifics (e.g., "Mr.", "Mrs.", "Dr.") that might have a period but are not the end of a sentence. - Negative lookbehind is also used to exclude some abbreviations (e.g., "Dept.", "Univ.", "et al.") that might have a period but are not the end of a sentence. - Positive lookbehind is used to match a whitespace character following a terminal punctuation mark (., !, ?, or :). """ punct_regex = re.compile(r"(?<=[.!?;:])(?:(?<!Prof\.)|(?<!Dept\.)|(?<!Univ\.)|(?<!et\sal\.))(?<!\w\.\w.)(?<![A-Z][a-z]\.)\s") return re.split(punct_regex, line)
测试用例
class TestSplitStringByPunctuation(object): def test_split_string_by_punctuation_1(self): # Test case 1 text1 = "I am studying at Univ. of California, Dept. of Computer Science. The research team includes " \ "Prof. Smith, Dr. Johnson, and Ms. Adams et al. so we are working on a new project." result1 = split_string_by_punctuation(text1) assert result1 == ['I am studying at Univ. of California, Dept. of Computer Science.', 'The research team includes Prof. Smith, Dr. Johnson, and Ms. Adams et al. ' 'so we are working on a new project.'], "Test case 1 failed" def test_split_string_by_punctuation_2(self): # Test case 2 text2 = "This is a city in U.S.A.. This is i.e. one! What about this e.g. one? " \ "Finally, here's the last one:" result2 = split_string_by_punctuation(text2) assert result2 == ['This is a city in U.S.A..', 'This is i.e. one!', 'What about this e.g. one?', "Finally, here's the last one:"], "Test case 2 failed" def test_split_string_by_punctuation_3(self): # Test case 3 text3 = "This sentence contains no punctuation marks from Mr. Zhong, Dr. Lu and Mrs. Han It should return as a single element list" result3 = split_string_by_punctuation(text3) assert result3 == [ 'This sentence contains no punctuation marks from Mr. Zhong, Dr. Lu and Mrs. Han It should return ' 'as a single element list'], "Test case 3 failed"
问题现象
测试未通过,以测试用例1为例,实际结果错误地在缩写处拆分:
['I am studying at Univ.', 'of California, Dept.', 'of Computer Science.', 'The research team includes Prof.', 'Smith, Dr. Johnson, and Ms. Adams et al.', 'so we are working on a new project.']
不符合预期的仅在句子结束标点后拆分的结果。
问题分析与解决方案
原正则表达式逻辑错误:使用(?:(?<!Prof\.)|(?<!Dept\.)|...)这种或逻辑的负向断言,意味着只要不满足其中一个条件就会匹配,完全搞反了需求——我们需要的是当标点属于这些缩写时不匹配拆分位置,也就是要同时排除所有缩写的情况,应该用多个负向断言叠加(而非或)。
另外,et\sal\.的写法错误,空格不需要转义,应改为et al\.。
修正后的函数:
import re def split_string_by_punctuation(line: str) -> list[str]: """ Splits a given string into a list of strings using terminal punctuation marks (., !, ?, or :) as delimiters. This function utilizes regular expression patterns to ensure that abbreviations, honorifics, and certain special cases are not considered as sentence delimiters. Args: line (str): The input string to be split into sentences. Returns: list: A list of strings representing the sentences obtained after splitting the input string. Notes: - Negative lookbehind is used to exclude abbreviations (e.g., "e.g.", "i.e.", "U.S.A."), which might have a period but are not the end of a sentence. - Negative lookbehind is also used to exclude honorifics (e.g., "Mr.", "Mrs.", "Dr.") that might have a period but are not the end of a sentence. - Negative lookbehind is also used to exclude some abbreviations (e.g., "Dept.", "Univ.", "et al.") that might have a period but are not the end of a sentence. - Positive lookbehind is used to match a whitespace character following a terminal punctuation mark (., !, ?, or :). """ # 修正后的正则:用多个负向断言叠加,排除所有缩写/头衔的情况 punct_regex = re.compile( r"(?<=[.!?;:])" r"(?<!Univ\.)(?<!Dept\.)(?<!Prof\.)(?<!et al\.)" r"(?<!Mr\.)(?<!Mrs\.)(?<!Ms\.)(?<!Dr\.)" r"(?<!e\.g\.)(?<!i\.e\.)(?<!\w\.\w\.)" r"\s" ) return re.split(punct_regex, line)
验证结果
修正后运行测试用例:
- 测试用例1:正确得到2个句子,不会在
Univ.、Dept.、Prof.、et al.处拆分 - 测试用例2:正确拆分出4个句子,保留
U.S.A.、i.e.、e.g.中的点号 - 测试用例3:返回单个元素列表,符合预期
内容的提问来源于stack exchange,提问作者Chloe
相关产品推荐
相关产品推荐

