如何移除从文件读取行后生成的Python列表中的空字符串?
解决方法:清理列表中的空字符串与空白内容
嘿,我来帮你搞定这个文本清理的问题!你从文件读入的列表里既有空字符串'',还有带首尾空白(甚至特殊字符)的行,这里有几种简洁高效的Python处理方案:
方案一:列表推导式(最直观)
首先建议读取文件时就处理UTF-8 BOM头(就是你示例里的),避免后续麻烦,然后用列表推导式一次性完成过滤和清理:
# 读取文件,用utf-8-sig自动去除BOM头 with open('data\\wonderland.txt', encoding='utf-8-sig') as file: lines_list = file.read().splitlines() # 过滤空字符串+清理每行首尾空白 cleaned_lines = [line.strip() for line in lines_list if line.strip()]
原理说明:
line.strip():会移除每行首尾的所有空白字符(空格、换行符、制表符等),同时如果有残留的BOM字符也会被去掉;if line.strip():确保只有清理后非空的行才会被保留,直接过滤掉原列表里的''和那些全是空白的行(比如' ')。
方案二:用filter+map组合(更简洁)
如果你喜欢函数式风格,可以用map做统一清理,再用filter过滤空内容:
cleaned_lines = list(filter(lambda line: line.strip(), map(str.strip, lines_list)))
或者结合文件读取一步到位:
with open('data\\wonderland.txt', encoding='utf-8-sig') as file: cleaned_lines = list(filter(lambda line: line.strip(), map(str.strip, file.read().splitlines())))
效果验证
用你给出的示例列表测试:
原列表片段:
[' Down the Rabbit-Hole','','','','Alice was beginning to get very tired of sitting by her sister on the','','bank, and of having nothing to do: once or twice she had peeped into the','',...]
处理后会得到:
['Down the Rabbit-Hole', 'Alice was beginning to get very tired of sitting by her sister on the', 'bank, and of having nothing to do: once or twice she had peeped into the', ...]
内容的提问来源于stack exchange,提问作者Houmes
相关产品推荐
相关产品推荐

