You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何过滤Pandas DataFrame中含'off-on'序列的元组列表?

移除Pandas DataFrame中含'off-on'相邻序列的元组列表

问题背景

需要处理如下Pandas DataFrame,过滤掉列l中包含'off'元组紧跟'on'元组的行:

import pandas as pd
import numpy as np

array = np.array([[1, [('on',1),('off',1),('off',1),('on',1)]], [2,[('off',1),('on',1),('on',1),('off',1)]]])
index_values = ['first', 'second']
column_values = ['id', 'l']
df = pd.DataFrame(data = array, 
                  index = index_values, 
                  columns = column_values)

原尝试代码因逻辑错误导致updated_col为空:

updated_col = []
for d in df['l'] : 
    for index, value in enumerate(d) : 
        if len(value) == index : 
            break 
        elif value[index] == 'off' and value[index + 1] == 'on' : 
            updated_col.append(value)

错误原因:混淆了元组元素与列表索引——value是单个元组(如('on',1)),value[index]取的是元组内的元素,而非列表中的下一个元组,逻辑完全偏离需求。

解决方案

方法1:使用lambda + apply过滤行

核心思路:对每个元组列表,检查是否存在相邻的('off', x)后跟('on', y),若不存在则保留该行。

# 定义判断函数:检查列表中是否存在off-on相邻序列
def has_off_on_sequence(lst):
    for i in range(len(lst)-1):
        if lst[i][0] == 'off' and lst[i+1][0] == 'on':
            return True
    return False

# 过滤掉包含该序列的行
filtered_df = df[~df['l'].apply(lambda x: has_off_on_sequence(x))]

如果想直接用lambda简写(无需单独定义函数):

filtered_df = df[~df['l'].apply(lambda lst: any(lst[i][0] == 'off' and lst[i+1][0] == 'on' for i in range(len(lst)-1)))]

方法2:修正pairwise函数后使用

你提供的pairwise函数用itertools.combinations会生成所有两两组合(非相邻),不符合“紧跟”的要求,需改成生成相邻元素对的版本:

import itertools

# 修正后的pairwise:生成相邻元素对
def pairwise(x):
    # Python3.10+ 可直接用 itertools.pairwise(x)
    return zip(x, x[1:])

# 过滤逻辑
filtered_df = df[~df['l'].apply(lambda lst: any(p[0][0] == 'off' and p[1][0] == 'on' for p in pairwise(lst)))]

结果说明

运行上述代码后,会过滤掉所有列l中存在off紧跟on的行。针对你的示例DataFrame,两行都包含该序列,所以最终filtered_df为空;若有行的元组列表无该序列,则会被保留。

内容的提问来源于stack exchange,提问作者blue-sky

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.17 18:01:18