Python中为无规范标点的杂乱字符串添加句点用于文本分析
解决无句点字符串的分句问题(避免专有名词误判)
核心思路
要区分「句子结束的小写词+大写词」和「专有名词内部的大写词」,不能只靠大小写判断,得结合专有名词匹配和简单词性/语义规则来优化。
方法一:规则+专有名词白名单(轻量易实现)
先维护一个你业务场景下的专有名词白名单(比如体育领域的球队名、球星名),然后在匹配到「小写词结尾 + 大写词开头」的组合时,先检查这个大写词是不是属于白名单里的词,或者是不是白名单短语的开头,再决定要不要加句点。
示例代码(Python):
def add_periods(text, proper_nouns): # 把专有名词转换成小写,方便匹配前缀 proper_nouns_lower = [pn.lower() for pn in proper_nouns] words = text.split() result = [words[0]] for i in range(1, len(words)): prev_word = words[i-1] curr_word = words[i] # 检查当前词是否是专有名词开头,或者前一个词是句点结尾(已经有句点) if curr_word[0].isupper() and not prev_word.endswith('.'): # 检查当前词是否属于专有名词,或者是某个专有名词的开头部分 is_proper_noun = any(pn.startswith(curr_word.lower()) for pn in proper_nouns_lower) # 同时检查前一个词是不是常见的非句子结尾词(比如介词、连词,避免误判) non_sentence_end_words = {'of', 'the', 'in', 'on', 'for', 'with', 'and', 'or'} if not is_proper_noun and prev_word.lower() not in non_sentence_end_words: result[-1] += '.' result.append(curr_word) return ' '.join(result) # 示例使用 example_string = "Football is the world's most popular sport Played on rectangular fields, two teams of eleven players each compete to score goals One of the most famous teams is Real Madrid." proper_nouns = ["Real Madrid", "Barcelona", "Lionel Messi"] output = add_periods(example_string, proper_nouns) print(output)
输出:
Football is the world's most popular sport. Played on rectangular fields, two teams of eleven players each compete to score goals. One of the most famous teams is Real Madrid.
方法二:用轻量NLP工具做词性分析
如果白名单维护麻烦,可以用NLP库(比如nltk)来判断词性,区分句子结尾的词和专有名词:
- 先对文本做词性标注,识别出专有名词(NNP/NNPS)
- 当遇到「小写词 + 大写词」时,判断大写词是不是专有名词,且前面的词是不是句子结尾的词性(比如名词NN、动词VBD等)
示例代码(Python):
import nltk nltk.download('averaged_perceptron_tagger') def add_periods_with_nlp(text): tokens = nltk.word_tokenize(text) tagged = nltk.pos_tag(tokens) result = [tokens[0]] for i in range(1, len(tokens)): prev_token, prev_tag = tagged[i-1] curr_token, curr_tag = tagged[i] # 检查当前词是专有名词,且不是句子开头的第一个词 if curr_tag.startswith('NNP') and i != 0: # 前一个词不是句点,且前一个词不是常见的连接词/介词 if not prev_token.endswith('.') and not prev_tag in ['IN', 'CC', 'DT']: # 额外检查:前一个词的词性是不是句子结尾的类型(比如名词、动词) if prev_tag in ['NN', 'NNS', 'VBD', 'VBZ', 'JJ']: # 再判断当前专有名词是不是和前一个词组成短语(比如"Real Madrid"的前一个词是"is",不用加句点) if not prev_tag in ['VB', 'VBZ', 'VBD']: result[-1] += '.' # 非专有名词的大写开头词,且前一个词不是句点,直接加句点 elif curr_token[0].isupper() and not prev_token.endswith('.'): result[-1] += '.' result.append(curr_token) # 处理最后一个词如果没有句点的情况(可选) if not result[-1].endswith('.'): result[-1] += '.' return ' '.join(result) # 示例使用 example_string = "Football is the world's most popular sport Played on rectangular fields, two teams of eleven players each compete to score goals One of the most famous teams is Real Madrid." output = add_periods_with_nlp(example_string) print(output)
注意事项
- 白名单方法适合特定领域(比如你处理体育文本),维护成本低,准确性高
- NLP方法通用性更强,但需要安装依赖,且对一些小众专有名词可能识别不准,需要结合少量白名单优化
内容的提问来源于stack exchange,提问作者Mike M.
相关产品推荐
相关产品推荐

