pandas筛选任意列含字符串的行 拆分可转换为整型的干净数据集
解决pandas含字符串异常行的拆分方案
核心逻辑是构造布尔掩码判断每行是否包含指定异常字符,再按掩码拆分数据集,无需手动指定列名,兼容任意列数的object类型DataFrame。
完整实现代码
import pandas as pd # 示例原始数据 df = pd.DataFrame({ 'col1':[1,2,'a',0,3], 'col2':[1,2,3,4,5], 'col3':[1,2,3,'45a5',4] }) # 1. 定义需要排查的异常字符集合,可按需扩展新增 invalid_chars = {'a', 'b'} # 2. 逐元素判断是否包含异常字符,按行聚合得到异常行掩码(只要有一列符合就标记为异常行) error_mask = df.applymap(lambda x: any(char in str(x) for char in invalid_chars)).any(axis=1) # 3. 拆分数据集 df_clean = df[~error_mask].reset_index(drop=True) df_error = df[error_mask].reset_index(drop=True)
运行结果验证
输出df_clean即为无异常字符的干净数据集,可直接转换为integer类型:
| col1 | col2 | col3 |
|---|---|---|
| 1 | 1 | 1 |
| 2 | 2 | 2 |
| 3 | 5 | 4 |
输出df_error即为包含异常字符的行集合:
| col1 | col2 | col3 |
|---|---|---|
| a | 3 | 3 |
| 0 | 4 | 45a5 |
注意事项
- 可灵活调整
invalid_chars集合的内容,自定义需要排查的异常字符 - 无需单独指定列名,自动适配任意列数的DataFrame
- 若不需要重置拆分后数据集的索引,可删除
reset_index(drop=True)代码段
内容的提问来源于stack exchange,提问作者fellowCoder
相关产品推荐
相关产品推荐

