You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何从pandas DataFrame提取"START"对应列的后续数值以得到目标k?

Pandas提取含"START"(大小写匹配)列的后续数值问题

需求:处理DataFrame时,提取所有存在"START"(大小写不敏感)的列中,该字符串所在行之后的数值数据,仅保留目标列,不要多余列。

原代码及问题

原代码

import pandas as pd

df = pd.DataFrame({'A': ['43', 23, 'Ndfg', 34, 0, 56],
               'B': ['5', 23, 'START', 89, 0, 4],
               'C': ['65', 7, 'dsfgA', 65, 47, 3],
               'D': ['65', 7, 'gfd', 3, 0, 7],
               'E': ['76', 7, 'Start', 5, 12, 1],
               'F': ['65', 7, 'sdfA', 5, 0, 4],
               'G': ['12', 7, 'START', 5, 8, 9],
               'H': ['89', 7, 'gfA', 5, 0, 8],
               'I': ['23', 7, 'sdfA', 5, 7, 23]})

k = []
for rw in range(df.shape[0]):
    for colm in range(df.shape[1]):
        if df.iloc[rw, colm] == 'START':
            row = rw + 1
            col = colm
            k.append(df.iloc[row:, col:])
            break

当前输出(包含多余列,且遗漏部分目标列)

[    B   C  D   E  F  G  H   I
 3  89  65  3   5  5  5  5   5
 4   0  47  0  12  0  8  0   7
 5   4   3  7   1  4  9  8  23]

期望输出

B   E  G
3  89   5  5
4   0  12  8
5   4   1  9

问题分析

  1. 多余列问题:原代码中df.iloc[row:, col:]会提取从匹配列开始的所有后续列,而非仅目标列。
  2. 大小写不匹配:仅判断== 'START',无法识别'Start'这类大小写不同的字符串。
  3. 遗漏目标列:循环中找到第一个匹配后执行break,仅处理了第一列(B列),未处理E、G列。

解决方案代码

import pandas as pd

df = pd.DataFrame({'A': ['43', 23, 'Ndfg', 34, 0, 56],
                   'B': ['5', 23, 'START', 89, 0, 4],
                   'C': ['65', 7, 'dsfgA', 65, 47, 3],
                   'D': ['65', 7, 'gfd', 3, 0, 7],
                   'E': ['76', 7, 'Start', 5, 12, 1],
                   'F': ['65', 7, 'sdfA', 5, 0, 4],
                   'G': ['12', 7, 'START', 5, 8, 9],
                   'H': ['89', 7, 'gfA', 5, 0, 8],
                   'I': ['23', 7, 'sdfA', 5, 7, 23]})

# 筛选目标列及对应起始行
target_cols = []
start_row_indices = []

for col_name in df.columns:
    # 大小写不敏感匹配,忽略空值
    has_start = df[col_name].str.contains('START', case=False, na=False)
    if has_start.any():
        target_cols.append(col_name)
        # 获取第一个匹配的行索引
        first_start_row = has_start.idxmax()
        start_row_indices.append(first_start_row)

# 提取每个目标列的后续数据并合并
result = pd.DataFrame()
for col, row_idx in zip(target_cols, start_row_indices):
    result[col] = df.loc[row_idx + 1:, col]

print(result)

代码说明

  • 使用str.contains('START', case=False)实现大小写不敏感匹配,na=False避免空值引发错误。
  • 遍历所有列,筛选出包含目标字符串的列,并记录每个列中第一个匹配的行索引。
  • 对每个目标列,提取匹配行之后的所有数据,合并成最终结果DataFrame,仅保留目标列。

内容的提问来源于stack exchange,提问作者shoggananna

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.21 23:30:47