如何用Python从文本文件提取指定列数据?Pandas调用失败求助
提取文本文件指定列的解决方案
原代码问题分析
你的代码无法运行主要有三个原因:
- 读取的文件是
.txt而非.csv,且数据用多空格分隔,不是逗号 - 未处理开头的注释行(以
#开头的行),导致pandas把注释内容当成数据读取 - 未指定要提取的目标列,且行范围选择错误
方案一:使用Pandas处理(推荐)
Pandas能高效处理带表头的结构化数据,适合批量文件操作。
单个文件处理
import pandas as pd # 读取txt文件,自动跳过注释行,处理多空格分隔 df = pd.read_csv( "input.txt", comment="#", # 跳过所有以#开头的行 sep=r"\s+", # 匹配任意数量的空格作为分隔符 engine="python" # 确保正则分隔符生效 ) # 提取目标列(D(MN/WN)),也可用iloc[:,2]按索引取第3列(从0开始计数) target_data = df["D(MN/WN)"] # 保存到新文件,每行一个数值 target_data.to_csv("output.txt", index=False, header=False)
批量处理多个txt文件
import pandas as pd import glob # 获取当前目录下所有txt文件路径 all_txt_files = glob.glob("*.txt") for file_path in all_txt_files: # 生成对应输出文件名(比如input.txt → input_output.txt) output_path = f"{file_path.split('.')[0]}_output.txt" # 读取文件 df = pd.read_csv( file_path, comment="#", sep=r"\s+", engine="python" ) # 保存目标列数据 df["D(MN/WN)"].to_csv(output_path, index=False, header=False)
方案二:纯Python处理(无需第三方库)
如果不想依赖Pandas,可直接用Python内置文件操作实现:
单个文件处理
with open("input.txt", "r") as input_file, open("output.txt", "w") as output_file: for line in input_file: stripped_line = line.strip() # 跳过空行和注释行 if not stripped_line or stripped_line.startswith("#"): continue # 按空格分割行内容 line_parts = stripped_line.split() try: # 验证第三项是数值,确认是数据行而非表头 float(line_parts[2]) # 写入目标值 output_file.write(f"{line_parts[2]}\n") except (ValueError, IndexError): # 跳过表头行或格式错误的行 continue
批量处理多个txt文件
import glob for file_path in glob.glob("*.txt"): output_path = f"{file_path.split('.')[0]}_output.txt" with open(file_path, "r") as input_file, open(output_path, "w") as output_file: for line in input_file: stripped_line = line.strip() if not stripped_line or stripped_line.startswith("#"): continue line_parts = stripped_line.split() try: float(line_parts[2]) output_file.write(f"{line_parts[2]}\n") except (ValueError, IndexError): continue
内容的提问来源于stack exchange,提问作者pro
相关产品推荐
相关产品推荐

