You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Polars数据框验证列表列连续相同元素最多连续出现两次

检查Polars DataFrame中列表的连续元素重复次数不超过2次

需求:验证DataFrame的list列中,每个整数列表的连续相同整数出现次数最多为2次,若存在连续3次及以上的重复元素,则触发断言失败。

解决方案

方法1:自定义函数逐行处理(直观易读)

通过自定义函数计算每个列表的最大连续重复次数,再批量验证:

import polars as pl

def get_max_consecutive(lst):
    if not lst:
        return 0
    current_count = 1
    max_count = 1
    for idx in range(1, len(lst)):
        if lst[idx] == lst[idx-1]:
            current_count += 1
            max_count = max(max_count, current_count)
        else:
            current_count = 1
    return max_count

# 构造示例DataFrame
row1 = [0, 1, -1, -1, 1, 1, -1, 0]
row2 = [1, -1, -1, -1, 0, 0, 1, -1]
df = pl.DataFrame({"list": [row1, row2]})

# 计算每行的最大连续重复次数
df = df.with_columns(
    max_consecutive = pl.col("list").map_elements(get_max_consecutive, return_dtype=pl.Int64)
)

# 执行断言验证
assert (df["max_consecutive"] <= 2).all(), "存在连续重复次数超过2次的列表"

执行后会触发断言错误,因为row2中的-1连续出现了3次,对应的max_consecutive值为3,不符合要求。

方法2:Polars原生函数(高效适合大数据量)

利用Polars内置的list.rle_lengths()函数直接获取连续元素的长度列表,再取最大值,避免Python层面的循环:

import polars as pl

row1 = [0, 1, -1, -1, 1, 1, -1, 0]
row2 = [1, -1, -1, -1, 0, 0, 1, -1]
df = pl.DataFrame({"list": [row1, row2]})

# 用原生函数计算最大连续重复次数
df = df.with_columns(
    max_consecutive = pl.col("list").list.rle_lengths().list.max()
)

# 执行断言验证
assert (df["max_consecutive"] <= 2).all(), "存在连续重复次数超过2次的列表"

list.rle_lengths()会将原列表转换为连续元素的长度列表,比如row2的列表会被转换为[1, 3, 2, 1, 1],最大值为3,因此触发断言失败。

说明

  • 方法1适合小数据量或需要自定义复杂逻辑的场景,代码直观易懂;
  • 方法2基于Polars原生实现,性能更优,适合处理大规模数据集。

内容的提问来源于stack exchange,提问作者Andi

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.19 03:42:17