Python 3如何移除字符串非空行前导空白以便转换为列表
Python 3 非空行前导空白移除方案
问题说明
作业要求移除输出文本中非换行内容行的所有前导空白,保留原有空行结构,方便后续字符串转列表;规则允许使用正则,禁止引入xml相关模块。
当前程序输出的非空内容行存在前导缩进,原始输出、预期输出分别如下:
原始带缩进输出
Belgian Waffles $5.95 Two of our famous Belgian Waffles with plenty of real maple syrup 650 Strawberry Belgian Waffles $7.95 Light Belgian waffles covered with strawberries and whipped cream 900 Berry-Berry Belgian Waffles $8.95 Light Belgian waffles covered with an assortment of fresh berries and whipped cream 900 French Toast $4.50 Thick slices made from our homemade sourdough bread 600 Homestyle Breakfast $6.95 Two eggs, bacon or sausage, toast, and our ever-popular hash browns 950 (end of output)
预期处理后输出
Belgian Waffles $5.95 Two of our famous Belgian Waffles with plenty of real maple syrup 650 Strawberry Belgian Waffles $7.95 Light Belgian waffles covered with strawberries and whipped cream 900 Berry-Berry Belgian Waffles $8.95 Light Belgian waffles covered with an assortment of fresh berries and whipped cream 900 French Toast $4.50 Thick slices made from our homemade sourdough bread 600 Homestyle Breakfast $6.95 Two eggs, bacon or sausage, toast, and our ever-popular hash browns 950 (end of output)
现有待修改代码
import os import re def get_filename(): print("Enter the name of the file: ") filename = input() return filename def read_file(filename): if os.path.exists(filename): with open(filename, "r") as file: full_text = file.read() return full_text else: print("This file does not exist") def get_tags(full_text): tags = re.findall('<.*?>', full_text) for tag in tags: full_text = full_text.replace(tag, '') return tags def get_text(text): tags = re.findall('<.*?>', text) for tag in tags: text = text.replace(tag, '') text = text.strip() return text def display_output(text): print(text) def main(): filename = get_filename() full_text = read_file(filename) tags = get_tags(full_text) text = get_text(full_text) display_output(text) main()
问题根因
现有get_text函数中对整个文本直接调用strip(),仅能移除整个字符串首尾的空白字符,既无法处理中间每一行的前导缩进,还会破坏原有的空行分段结构。
修复方案
方案1:正则实现(符合作业允许使用RegEx的规则)
使用正则匹配每一行开头的所有连续空白字符,替换为空即可。需要给正则加re.MULTILINE标志,让^元字符匹配每一行的开头,而非仅匹配整个字符串的起始位置,空行因为没有非空白内容不会被误处理,原有换行结构会完整保留。
仅需修改get_text函数即可:
def get_text(text): # 移除所有类HTML标签 text = re.sub(r'<.*?>', '', text) # 逐行移除前导空白,re.M标志开启多行匹配模式 text = re.sub(r'^\s+', '', text, flags=re.MULTILINE) return text
注:原代码中
get_tags和get_text重复做了标签匹配替换的逻辑,这里直接用re.sub一次性替换所有标签,比循环replace效率更高。
方案2:字符串内置方法实现(无需正则)
将文本按换行符拆分为行列表,对每一行单独调用lstrip()移除左侧空白,处理完成后再用换行符拼接回完整字符串,空行调用lstrip()后仍为空字符串,拼接后原有空行结构完全保留:
def get_text(text): # 移除所有类HTML标签 text = re.sub(r'<.*?>', '', text) # 逐行处理左空白 line_list = text.split('\n') text = '\n'.join([line.lstrip() for line in line_list]) return text
后续转列表提示
处理完成的文本直接调用split('\n')即可转为列表,如果不需要保留空行元素,可在列表推导中加判断过滤:
# 保留空行的列表 full_list = processed_text.split('\n') # 过滤空行的列表 non_empty_list = [line for line in processed_text.split('\n') if line]
内容的提问来源于stack exchange,提问作者maxmansupa
相关产品推荐
相关产品推荐

