You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用Polars实现数据集字符串去空格的有效性测试?

用Polars替代Python循环验证字符串去空格函数的测试逻辑

问题描述

我已经用Polars开发了一个数据集字符串去空格函数,现在需要编写测试验证去空格操作是否成功。目前我有如下基于Python循环的测试逻辑代码,请问如何改用Polars实现该测试?

原测试代码:

def test_strip():
    df = pd.DataFrame({
        'ID': [1, 1, 1, 1, 1],
        'Entity': ['Entity 1 ', 'Entity 2', 'Entity 3', 'Entity 4', 'Entity 5'],
        'Table': ['Table 1', ' Table 2', 'Table 3', 'Table 4', None],
        'Local': ['Local 1', 'Local 2 ', None, 'Local 4', 'Local 5'],
        'Global': ['Global 1', ' Global 2', 'Global 3', None, ' Global 5'],
        'mandatory': ['M', 'M', 'M', 'CM ', 'M']
    })
    job = first_job(
        config=test_config,
        copying_list=copying,
    )
    result = job.run(df)
    df_clean, *_ = result

    for column in df_clean.columns:
        for value in df_clean[column]:
            if isinstance(value, str) and (value.startswith(" ") or value.endswith(" ")):
                raise AssertionError(f"Strip failed for column '{column}'")

解决方案

方式一:逐列验证(贴近原逻辑)

用Polars的矢量化操作替代Python循环,针对每个字符串列检查是否存在首尾空格,效率更高且符合Polars的API风格:

def test_strip():
    # 若数据流程基于Polars,可直接用pl.DataFrame替代pd.DataFrame
    df = pl.DataFrame({
        'ID': [1, 1, 1, 1, 1],
        'Entity': ['Entity 1 ', 'Entity 2', 'Entity 3', 'Entity 4', 'Entity 5'],
        'Table': ['Table 1', ' Table 2', 'Table 3', 'Table 4', None],
        'Local': ['Local 1', 'Local 2 ', None, 'Local 4', 'Local 5'],
        'Global': ['Global 1', ' Global 2', 'Global 3', None, ' Global 5'],
        'mandatory': ['M', 'M', 'M', 'CM ', 'M']
    })
    job = first_job(
        config=test_config,
        copying_list=copying,
    )
    result = job.run(df)
    df_clean, *_ = result

    # 仅筛选字符串类型的列,避免干扰非字符串列
    string_columns = df_clean.select(pl.Utf8).columns
    
    for col in string_columns:
        # 检查该列是否存在首尾带空格的字符串
        has_invalid_value = df_clean.select(
            pl.col(col).str.starts_with(" ").or(pl.col(col).str.ends_with(" "))
        ).any().item()
        
        assert not has_invalid_value, f"Strip failed for column '{col}'"

方式二:全局验证(快速定位问题行)

如果需要一次性找出所有存在问题的行,可使用以下方式,测试失败时直接输出问题行便于排查:

def test_strip():
    df = pl.DataFrame({
        'ID': [1, 1, 1, 1, 1],
        'Entity': ['Entity 1 ', 'Entity 2', 'Entity 3', 'Entity 4', 'Entity 5'],
        'Table': ['Table 1', ' Table 2', 'Table 3', 'Table 4', None],
        'Local': ['Local 1', 'Local 2 ', None, 'Local 4', 'Local 5'],
        'Global': ['Global 1', ' Global 2', 'Global 3', None, ' Global 5'],
        'mandatory': ['M', 'M', 'M', 'CM ', 'M']
    })
    job = first_job(
        config=test_config,
        copying_list=copying,
    )
    result = job.run(df)
    df_clean, *_ = result

    # 筛选出任意字符串列存在首尾空格的行
    problematic_rows = df_clean.filter(
        pl.any_horizontal(
            pl.col(pl.Utf8).str.starts_with(" ").or(pl.col(pl.Utf8).str.ends_with(" "))
        )
    )

    # 断言无问题行,失败时打印问题行详情
    assert problematic_rows.height == 0, f"Strip failed for rows:\n{problematic_rows}"

关键说明

  • pl.col(pl.Utf8)精准定位字符串类型列,避免对数值、None等非字符串类型做无效检查
  • str.starts_with()和str.ends_with()是Polars内置的字符串处理方法,支持矢量化操作,比Python循环高效得多
  • any()用于判断列中是否存在符合条件的值,any_horizontal()用于判断行内任意列是否符合条件

内容的提问来源于stack exchange,提问作者Horseman

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.01 09:45:27