Polars中能否在with_columns语句中通过布尔判断实现分类型转换并哈希?
Polars中能否在with_columns语句中通过布尔判断实现分类型转换并哈希?
当然可以实现啦!我来帮你搞定这个问题~
首先得理清你当前遇到的问题:你把字符串列stris也强行转成了pl.List(pl.Categorical),相当于每个独立的字符串都被包装成了单元素列表,Polars对这种特殊结构的哈希处理出现了重复,这才导致所有stris_hashed的值都一模一样,完全不符合你的预期。
要实现“根据列类型自动选择转换方式”的需求,你可以用Polars的when/then/otherwise条件分支逻辑,在同一个with_columns调用里完成不同列的差异化处理。具体代码如下:
list_of_lists = [ ['base', 'base.current base', 'base.current base.inventories - total', 'ABCD'], ['base', 'base.current base', 'base.current base.inventories - total', 'DEFG'], ['base', 'base.current base', 'base.current base.inventories - total', 'ABCD'], ['base', 'base.current base', 'base.current base.inventories - total', 'HIJK'] ] list_of_strings = ['(bobbyJoe460)', 'bobby, Joe (xx866e)', '137642039575', 'mamamia'] pl_df_1 = pl.DataFrame({'lists': list_of_lists,'stris':list_of_strings}, strict=False) # 修正后的代码 result_df = pl_df_1.with_columns( pl.col(['lists', 'stris']) # 判断当前列是否为List类型 .when(pl.col.dtype.is_(pl.List)) # 如果是List列,转成List(Categorical) .then(pl.col.cast(pl.List(pl.Categorical))) # 如果是其他类型(这里是字符串列),直接转成Categorical .otherwise(pl.col.cast(pl.Categorical)) # 执行哈希操作 .hash(seed=140) # 给哈希后的列名加后缀 .name.suffix('_hashed') ) print(result_df)
运行这段代码后,你会看到:
- 列表列
lists中内容相同的行,哈希值依然保持一致(比如第1行和第3行的lists_hashed相同) - 字符串列
stris的每个值都有了唯一的哈希值,不会再出现全相同的情况
这个逻辑的核心是用pl.col.dtype.is_(pl.List)来检测列的类型,然后针对性地应用不同的转换规则,完美适配你同时处理列表列和字符串列的需求。
备注:内容来源于stack exchange,提问作者MikeB2019x
相关产品推荐
相关产品推荐

