如何将段落拆分为单行句子并写入TXT?代码报错求助
问题分析与解决方案
首先,你的代码触发IndexError的原因很明确:
- 当你用
for char in document遍历字符串时,char是单个字符(比如第一个循环里是'H',第二个是'e'),你再用char[pos]去访问它的索引,当pos大于0时,单个字符的索引范围只有0,自然就超出范围了。 - 另外,Python中的字符串是不可变对象,你不能直接通过
char[pos] = "\n"这种方式修改字符内容。
下面给你两种简单可行的实现方式:
方法1:使用字符串分割与连接(最简洁)
利用split('.')按句号分割字符串,然后过滤掉空内容,再用换行符连接起来:
document = "Hello World. Goodbye World" def sentence_separator(document): # 按句号分割,去掉分割后可能的空字符串,再用换行符连接 sentences = [s.strip() for s in document.split('.') if s.strip()] return '\n'.join(sentences) print(sentence_separator(document))
输出结果:
Hello World Goodbye World
方法2:手动遍历构建新字符串
如果你想手动实现遍历逻辑,可以逐个字符处理,遇到句号时添加换行符:
document = "Hello World. Goodbye World" def sentence_separator(document): result = [] for char in document: if char == '.': # 遇到句号时,添加换行符,而不是直接替换句号 result.append('\n') else: result.append(char) # 把列表转成字符串,再处理可能的多余空格(比如句号后的空格) return ''.join(result).replace('\n ', '\n') print(sentence_separator(document))
这个方法也能得到同样的输出,而且更贴近你原本的遍历思路,只是避开了字符串不可变和索引错误的问题。
另外,如果你需要把结果写入TXT文件,可以在函数里加上文件写入逻辑,比如:
def sentence_separator(document, output_file): sentences = [s.strip() for s in document.split('.') if s.strip()] content = '\n'.join(sentences) with open(output_file, 'w') as f: f.write(content) return content # 调用示例 sentence_separator(document, 'output.txt')
这样就能直接生成每个句子单独一行的TXT文件了。
内容的提问来源于stack exchange,提问作者Liviu Iosif
相关产品推荐
相关产品推荐

