Python处理大文件时无法去除^@格式错误,求解决方案
问题描述
处理一个3万多行的数据文件时,遇到一行格式错误的数据:
^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@ 3.352388740E-05 1.956399032E+00 5.330716152E+09
正常数据行格式如下:
3.352202475E-05 1.956847538E+00 5.329098588E+09 3.351643682E-05 1.957293472E+00 5.327447942E+09 3.351271152E-05 1.957742024E+00 5.325773881E+09
后续将拆分后的行转换为float类型时失败,原因是无法将^@转换为浮点数。尝试过以下代码去除该字符或删除整行,但^@始终无法去除:
def clean_file(new_file): new_file_path = Path(new_file) lines = [] with new_file_path.open('r') as f: lines = f.readlines() with new_file_path.open('w') as file: for number, line in enumerate(lines): if number not in [7] and not line.startswith('^@'): file.write(line)
注:第7行开始是数据行,之前是表头。也用过line.replace('^@',''),同样无效。
解决方案
方法一:删除包含无效字符的行
^@不是字面的^加@,而是ASCII空字符(对应转义字符\x00),所以之前的字符串匹配方法都不生效。可以直接检查行中是否包含空字符,包含则跳过写入:
from pathlib import Path def clean_file(new_file): new_file_path = Path(new_file) with new_file_path.open('r') as f: lines = f.readlines() with new_file_path.open('w') as file: for line in lines: if '\x00' not in line: file.write(line)
如果需要保留前7行的表头,只过滤数据行:
from pathlib import Path def clean_file(new_file): new_file_path = Path(new_file) with new_file_path.open('r') as f: lines = f.readlines() with new_file_path.open('w') as file: for idx, line in enumerate(lines): if idx < 7 or '\x00' not in line: file.write(line)
方法二:去除行中的空字符
如果不想删除整行,只需去掉空字符,用replace('\x00', '')即可:
from pathlib import Path def clean_file(new_file): new_file_path = Path(new_file) with new_file_path.open('r') as f: lines = f.readlines() with new_file_path.open('w') as file: for line in lines: cleaned_line = line.replace('\x00', '') file.write(cleaned_line)
处理后错误行会变为正常数值行,后续转换float类型不会报错。
内容的提问来源于stack exchange,提问作者MOB100
相关产品推荐
相关产品推荐

