字符串中句子出现次数统计:单行与跨行两种实现需求
统计目标句子出现次数的两种实现方案
好的,我给你准备了两种实用的Python实现方案,分别对应单行完整匹配和跨行拆分匹配的需求:
1. 单行完整匹配版本
这个版本适用于统计文件中完整出现在单行的目标句子次数,逻辑很直接:逐行读取文件内容,检查每行里目标句子的出现次数并累加。
def count_full_line_occurrences(file_path, target_sentence): occurrence_count = 0 # 打开文件逐行读取 with open(file_path, 'r', encoding='utf-8') as file: for line in file: # 统计当前行中目标句子的出现次数(支持一行多次匹配) occurrence_count += line.count(target_sentence) return occurrence_count # 调用示例 target = '"My name is scott"' # 你的目标句子 file_path = 'your_target_file.txt' # 替换成你的文件路径 result = count_full_line_occurrences(file_path, target) print(f"单行完整匹配的次数:{result}")
2. 跨行拆分匹配版本
这个版本可以处理目标句子被拆分成多个字符串(比如"My name is " + "scott")且可能跨行的情况。核心思路是先把文件内容整体读入,再把所有拼接形式的字符串合并成完整句子,最后统计次数。
import re def count_split_line_occurrences(file_path, target_sentence): # 读取整个文件内容 with open(file_path, 'r', encoding='utf-8') as file: full_content = file.read() # 用正则替换所有字符串拼接的部分(包括跨行的" + ") # 把类似"xxx" + "yyy"或者跨行的"xxx"\n+\n"yyy"合并成"xxxyyy" processed_content = re.sub(r'"[\s\n]*\+[\s\n]*"', '', full_content) # 统计处理后内容中目标句子的出现次数 occurrence_count = processed_content.count(target_sentence) return occurrence_count # 调用示例 target = '"My name is scott"' file_path = 'your_target_file.txt' result = count_split_line_occurrences(file_path, target) print(f"跨行拆分匹配的次数:{result}")
补充说明
- 如果你需要更精确的匹配(比如避免部分匹配的干扰),可以把
count方法换成正则匹配的方式,比如用len(re.findall(re.escape(target_sentence), processed_content))来统计,这样能避免子串匹配的问题。 - 注意文件编码,如果你的文件不是UTF-8,需要调整
open函数里的encoding参数。
内容的提问来源于stack exchange,提问作者scott miles
相关产品推荐
相关产品推荐

