You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Pandas筛选制表符文本数据:保留对应下一行值为0的元素

问题描述

我有一个格式如下的txt文件:

0.0  0    5    6.31000
      5.29559    2.38176    0.51521    0.04454    0.00000
            0          0          0          0          2
  0.0  0    4    6.31000
      4.32454    1.77600    0.04454    0.00000
            0          0          0          2
  0.0  0    2    6.31000
      1.55590    0.00000
            0          0
  0.0  0    6    6.31000
      5.37285    4.39339    3.56905    0.83230    0.04454    0.00000
            0          0          0          0          0          2
  0.0  0    3    6.31000
      4.22062    1.60321    0.00000
            0          0          0

我希望创建一个DataFrame,步骤要求:

  • 删除第1、4、7、10等行(即每3行的第一行)
  • 保留中间行中,对应下一行(第三行)位置值为0的元素。比如保留5.29559,因为它对应下一行的位置值是0;不保留0.00000,因为对应下一行的位置值是2。

最终期望的DataFrame格式如下:

5.29559    2.38176    0.51521    0.04454    

  4.32454    1.77600    0.04454    

  1.55590    0.00000

  5.37285    4.39339    3.56905    0.83230    0.04454    

  4.22062    1.60321    0.00000

我尝试了以下代码,但未得到预期结果:

import pandas as pd

filename = r"data.txt"

# This returns a dataframe with a single column
df = pd.read_table(filename, header=None)
# remove first row
df = df.drop(index=0)                                                                                                                                                                                            df = df.drop(df.index[df.index % 3 == 0])
# remove spaces from beginning of all rows
df = df.applymap(lambda x: x.lstrip())
df.to_csv('file_with_quotes.txt', sep=' ', index=False)
with open('file_with_quotes.txt', 'r') as f1, open('file_without_quotes.txt', 'w') as f2:
    line_num = 0
    for line in f1:
        if line_num == 0:
            line_num += 1
            continue
        line = line.strip().replace('"', '')
        f2.write(line + '\n')
import os
#delete the file using os.remove()
os.remove('file_with_quotes.txt')
df_new = pd.read_table('file_without_quotes.txt', header=None)
df_new = df[0].str.split("  ", expand=True)

df_new = df_new.replace('', pd.np.nan)
# replace NaN values with empty string
df_new = df_new.fillna("")

# remove empty strings using applymap()
df_new = df_new.applymap(lambda x: x if x != "" else None)

(df_new.stack()
   .groupby(level=0)
   .apply(lambda df: df.reset_index(drop=True))
   .unstack())
解决方案

你的代码绕了太多弯路,且没有处理“根据下一行的0/2筛选中间行元素”这个核心需求。下面是更简洁且符合要求的实现:

import pandas as pd

# 读取文件,按任意数量空格分割列,自动跳过空值和行首空格
df = pd.read_csv("data.txt", sep=r"\s+", header=None, engine="python")

result_rows = []
# 每3行一组遍历处理
for i in range(0, len(df), 3):
    # 提取每组的第2行(数据行)和第3行(筛选条件行),去掉空值
    data_row = df.iloc[i+1].dropna().values
    filter_row = df.iloc[i+2].dropna().values
    
    # 只保留筛选条件为0的对应数据元素
    filtered_data = data_row[filter_row == 0]
    result_rows.append(filtered_data)

# 将结果转为DataFrame,缺失值填充为空字符串
final_df = pd.DataFrame(result_rows).fillna("")
print(final_df)

代码说明

  1. 读取文件:用sep=r"\s+"按任意数量空格分割列,自动处理行首空格和空值,直接得到结构化的DataFrame,无需手动处理字符串。
  2. 分组处理:按每3行一组遍历,每组提取数据行和筛选行。
  3. 筛选元素:利用布尔索引,精准保留筛选行中值为0的对应位置的数据元素。
  4. 生成结果:把筛选后的每行数据转为DataFrame,缺失值填充为空字符串,完全匹配你要的格式。

运行这段代码后,就能得到预期的结果。


内容的提问来源于stack exchange,提问作者Montana

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.25 00:04:58