You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

正则实现按标点拆分字符串时,如何避免误拆分缩写?

按标点拆分句子时正确保留缩写并避免错误拆分

需求

需要实现一个函数,按句号、感叹号、问号、分号、冒号拆分包含多句子的字符串,同时满足:

  • 不拆分缩写(如Univ.、Dept.、Prof.、et al.等)
  • 保留缩写中的点号,不能像之前参考的方案那样移除这类点号(比如把U.S.A.转为USA)

现有实现代码

import re

def split_string_by_punctuation(line: str) -> list[str]:
    """
    Splits a given string into a list of strings using terminal punctuation marks (., !, ?, or :) as delimiters.

    This function utilizes regular expression patterns to ensure that abbreviations, honorifics,
    and certain special cases are not considered as sentence delimiters.

    Args:
        line (str): The input string to be split into sentences.

    Returns:
        list: A list of strings representing the sentences obtained after splitting the input string.

    Notes:
        - Negative lookbehind is used to exclude abbreviations (e.g., "e.g.", "i.e.", "U.S.A."),
          which might have a period but are not the end of a sentence.
        - Negative lookbehind is also used to exclude honorifics (e.g., "Mr.", "Mrs.", "Dr.")
          that might have a period but are not the end of a sentence.
        - Negative lookbehind is also used to exclude some abbreviations (e.g., "Dept.", "Univ.", "et al.")
          that might have a period but are not the end of a sentence.
        - Positive lookbehind is used to match a whitespace character following a terminal
          punctuation mark (., !, ?, or :).
    """
    punct_regex = re.compile(r"(?<=[.!?;:])(?:(?<!Prof\.)|(?<!Dept\.)|(?<!Univ\.)|(?<!et\sal\.))(?<!\w\.\w.)(?<![A-Z][a-z]\.)\s")


    return re.split(punct_regex, line)

测试用例

class TestSplitStringByPunctuation(object):
    def test_split_string_by_punctuation_1(self):
        # Test case 1
        text1 = "I am studying at Univ. of California, Dept. of Computer Science. The research team includes " \
                "Prof. Smith, Dr. Johnson, and Ms. Adams et al. so we are working on a new project."
        result1 = split_string_by_punctuation(text1)
        assert result1 == ['I am studying at Univ. of California, Dept. of Computer Science.',
                           'The research team includes Prof. Smith, Dr. Johnson, and Ms. Adams et al. '
                           'so we are working on a new project.'], "Test case 1 failed"

    def test_split_string_by_punctuation_2(self):
        # Test case 2
        text2 = "This is a city in U.S.A.. This is i.e. one! What about this e.g. one? " \
                "Finally, here's the last one:"
        result2 = split_string_by_punctuation(text2)
        assert result2 == ['This is a city in U.S.A..', 'This is i.e. one!', 'What about this e.g. one?',
                           "Finally, here's the last one:"], "Test case 2 failed"

    def test_split_string_by_punctuation_3(self):
        # Test case 3
        text3 = "This sentence contains no punctuation marks from Mr. Zhong, Dr. Lu and Mrs. Han It should return as a single element list"
        result3 = split_string_by_punctuation(text3)
        assert result3 == [
            'This sentence contains no punctuation marks from Mr. Zhong, Dr. Lu and Mrs. Han It should return '
            'as a single element list'], "Test case 3 failed"

问题现象

测试未通过,以测试用例1为例,实际结果错误地在缩写处拆分:

['I am studying at Univ.',
'of California, Dept.',
'of Computer Science.',
'The research team includes Prof.',
'Smith, Dr. Johnson, and Ms. Adams et al.',
'so we are working on a new project.']

不符合预期的仅在句子结束标点后拆分的结果。

问题分析与解决方案

原正则表达式逻辑错误:使用(?:(?<!Prof\.)|(?<!Dept\.)|...)这种或逻辑的负向断言,意味着只要不满足其中一个条件就会匹配,完全搞反了需求——我们需要的是当标点属于这些缩写时不匹配拆分位置,也就是要同时排除所有缩写的情况,应该用多个负向断言叠加(而非或)。

另外,et\sal\.的写法错误,空格不需要转义,应改为et al\.。

修正后的函数:

import re

def split_string_by_punctuation(line: str) -> list[str]:
    """
    Splits a given string into a list of strings using terminal punctuation marks (., !, ?, or :) as delimiters.

    This function utilizes regular expression patterns to ensure that abbreviations, honorifics,
    and certain special cases are not considered as sentence delimiters.

    Args:
        line (str): The input string to be split into sentences.

    Returns:
        list: A list of strings representing the sentences obtained after splitting the input string.

    Notes:
        - Negative lookbehind is used to exclude abbreviations (e.g., "e.g.", "i.e.", "U.S.A."),
          which might have a period but are not the end of a sentence.
        - Negative lookbehind is also used to exclude honorifics (e.g., "Mr.", "Mrs.", "Dr.")
          that might have a period but are not the end of a sentence.
        - Negative lookbehind is also used to exclude some abbreviations (e.g., "Dept.", "Univ.", "et al.")
          that might have a period but are not the end of a sentence.
        - Positive lookbehind is used to match a whitespace character following a terminal
          punctuation mark (., !, ?, or :).
    """
    # 修正后的正则:用多个负向断言叠加,排除所有缩写/头衔的情况
    punct_regex = re.compile(
        r"(?<=[.!?;:])"
        r"(?<!Univ\.)(?<!Dept\.)(?<!Prof\.)(?<!et al\.)"
        r"(?<!Mr\.)(?<!Mrs\.)(?<!Ms\.)(?<!Dr\.)"
        r"(?<!e\.g\.)(?<!i\.e\.)(?<!\w\.\w\.)"
        r"\s"
    )
    return re.split(punct_regex, line)

验证结果

修正后运行测试用例:

  • 测试用例1:正确得到2个句子,不会在Univ.、Dept.、Prof.、et al.处拆分
  • 测试用例2:正确拆分出4个句子,保留U.S.A.、i.e.、e.g.中的点号
  • 测试用例3:返回单个元素列表,符合预期

内容的提问来源于stack exchange,提问作者Chloe

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.25 15:57:08