如何用Python从指定文本中提取完整有效句子?
实现思路
步骤拆解
- 文本预处理:将原始文本按换行符分割为单行,清理每行首尾的空白字符(多余空格、制表符等),同时过滤掉空行。
- 拆分前缀与内容:对每一行,以第一个冒号为分隔点拆分,分离出前缀标签(如
Hobby:、Something (where):)和后续内容,避免误拆分前缀里的特殊字符(比如括号内的内容)。 - 筛选完整句子:通过规则判断内容是否为完整句子:
- 基础规则:内容需以大写字母开头、以句号结尾;
- 补充规则:排除无主谓结构的片段(比如示例中的
To Everest mountain,仅为介词短语,无主语和谓语动词),可通过检查内容中是否包含常见谓语动词(如like、want、go等)辅助判断。
- 收集结果:将所有符合条件的完整句子收集起来,得到最终提取结果。
代码示例(Python)
document = "Hobby: I like going to the mountains.\n Something (where): To Everest mountain.\n\n The reason: I want to go because I like nature.\n Activities: I'd like to go hiking and admiring the beauty of the nature. " # 分割并清理行数据 processed_lines = [line.strip() for line in document.split('\n') if line.strip()] valid_sentences = [] for line in processed_lines: if ':' in line: # 按第一个冒号拆分,避免前缀含特殊字符时出错 _, content = line.split(':', 1) content = content.strip() # 验证完整句子的规则 if content and content[0].isupper() and content.endswith('.'): # 检查是否存在谓语动词,排除非完整片段 core_verbs = {'like', 'want', 'go', 'admire', 'hike'} if any(verb in content.lower() for verb in core_verbs): valid_sentences.append(content) print(valid_sentences) # 输出结果: # ['I like going to the mountains.', 'I want to go because I like nature.', "I'd like to go hiking and admiring the beauty of the nature."]
内容的提问来源于stack exchange,提问作者stephsmith
相关产品推荐
相关产品推荐

