如何修改Python脚本:忽略行尾分隔符并统计有效分隔符
问题:统计行内有效分隔符并忽略行尾分隔符
现有一段Python脚本,用于统计文本行中分隔符(, .)的数量,并将统计数添加至行首数字前缀后写入新文件。当前需求为忽略行尾的分隔符,仅统计行内有效分隔符数量。示例如下:
输入:
"1:2:3 ThisIs,Test,"
期望输出:"1:2:3:1 ThisIs,Test,"(行尾逗号不计入统计)
原脚本代码:
with open('123.txt', encoding='utf-8') as input_file: texts = input_file.readlines() delimiters = ',.' for text in texts: count = 0 # We count the number of characters from the list of delimiters in the given line for character in text: if character in delimiters: count += 1 # We add the number of characters from the delimiters list to the end of the line numbers text_parts = text.split(' ') text_parts[0] = text_parts[0] + ':' + str(count) new_text = ' '.join(text_parts) print(new_text, file=open(""+str("10000")+".txt", "a", encoding='utf-8'))
修改方案
核心思路:先定位行内最后一个非分隔符的位置,仅统计该位置之前的分隔符数量,同时优化文件操作逻辑。
修改后的完整代码:
with open('123.txt', encoding='utf-8') as input_file: texts = input_file.readlines() delimiters = ',.' # 一次性打开输出文件,避免循环内重复IO操作 with open("10000.txt", "w", encoding='utf-8') as output_file: for text in texts: # 移除行尾换行符,避免干扰行尾分隔符判断 stripped_line = text.rstrip('\n') count = 0 # 从行尾向前遍历,跳过所有连续的分隔符 end_pos = len(stripped_line) - 1 while end_pos >= 0 and stripped_line[end_pos] in delimiters: end_pos -= 1 # 统计有效范围内的分隔符数量 for char in stripped_line[:end_pos + 1]: if char in delimiters: count += 1 # 分割行首前缀与内容(仅分割1次,避免内容含空格被拆分) text_parts = text.split(' ', 1) if text_parts: text_parts[0] = f"{text_parts[0]}:{count}" new_text = ' '.join(text_parts) print(new_text, file=output_file)
关键优化点
- 用
rstrip('\n')仅移除换行符,保留行内其他空白字符 - 通过反向遍历定位有效内容的结束位置,精准跳过行尾所有连续分隔符
- 使用
split(' ', 1)避免内容中的空格被错误分割 - 一次性打开输出文件,提升代码运行效率
内容的提问来源于stack exchange,提问作者ZION BEAST
相关产品推荐
相关产品推荐

