如何使用Python移除文本中的连字符并保留原有格式
问题描述
需要编写Python函数,将带跨行连字符的文本转换为保留原有换行和缩进、且移除跨行连字符的格式:
输入文本:
text = '''hi guy- s how do i re- order this text so it doesn- t have any "-" ele- ments and it is still i this form'''
预期输出:
text = '''hi guys how do i reorder this text so it doesnt have any "-" elements and it is still i this form'''
之前尝试的代码未达到预期效果:
def uprav(text): zoz = "" text = text.split() print(len(text)) for i in text: if i[-1] == "-": nove_slovo = i[:-1] zoz = zoz + nove_slovo else: zoz = zoz + i + " " zoz.split() print(zoz)
问题分析
原代码的核心问题是:
- 使用
text.split()拆分文本时,会将所有空白字符(换行、多空格)统一拆分,彻底丢失了原有的换行结构和每行缩进信息。 - 处理逻辑仅拼接单词,无法保留原文本的行格式,最终输出是无结构的字符串。
解决方案
要保留原文本的换行和缩进,需按行处理,识别跨行连字符并拼接对应的单词,同时保留行结构。以下是实现代码:
def fix_hyphenated_lines(text): lines = text.split('\n') result = [] i = 0 while i < len(lines): current_line = lines[i] # 检查当前行是否以连字符结尾 if current_line.endswith('-'): # 去掉末尾的连字符 prefix = current_line[:-1] i += 1 if i >= len(lines): result.append(prefix) break next_line = lines[i] # 拆分下一行的前置空格、第一个单词和剩余内容 leading_spaces = len(next_line) - len(next_line.lstrip()) parts = next_line.split(maxsplit=1) if parts: first_word = parts[0] # 拼接前缀与第一个单词,作为新的行加入结果 result.append(prefix + first_word) # 处理下一行的剩余内容:如果还有内容,保留缩进继续循环 if len(parts) > 1: lines[i] = ' ' * leading_spaces + parts[1] else: i += 1 else: result.append(prefix) i += 1 else: result.append(current_line) i += 1 return '\n'.join(result)
测试验证
调用函数测试输入文本:
text = '''hi guy- s how do i re- order this text so it doesn- t have any "-" ele- ments and it is still i this form''' print(fix_hyphenated_lines(text))
输出结果与预期完全一致:
hi guys how do i reorder this text so it doesnt have any "-" elements and it is still i this form
内容的提问来源于stack exchange,提问作者Simona
相关产品推荐
相关产品推荐

