如何在出现连字符时拆分DataFrame并删除对应行
拆分含连字符的DataFrame并删除对应行
核心思路
先定位所有包含连字符-的行,再以这些行为分隔点,将原DataFrame切割成多个连续的、不含连字符的子DataFrame,同时丢弃含连字符的行。
代码实现
1. 读取数据
假设你的数据是分号分隔格式,先读取为pandas DataFrame:
import pandas as pd # 替换为你的数据路径或读取逻辑 df = pd.read_csv('your_data_file.csv', sep=';', header=0)
2. 定位含连字符的行
遍历每行,标记并提取包含连字符的行索引:
# 把每行转为字符串后检查是否存在连字符'-' has_hyphen = df.apply(lambda row: '-' in row.astype(str).values, axis=1) # 获取所有含连字符行的索引列表 hyphen_positions = df[has_hyphen].index.tolist()
3. 拆分DataFrame
根据定位到的索引切割原DataFrame,生成子DataFrame列表:
split_dfs = [] current_start = 0 for pos in hyphen_positions: # 提取从current_start到pos前一行的有效数据 if current_start < pos: split_dfs.append(df.iloc[current_start:pos].copy()) # 更新起始位置为当前含连字符行的下一行 current_start = pos + 1 # 处理最后一段未被切割的有效数据 if current_start < len(df): split_dfs.append(df.iloc[current_start:].copy())
结果说明
- 执行后,
split_dfs就是拆分后的所有子DataFrame列表,对应你示例中的预期结果(2个DataFrame) - 如果需要重置每个子DataFrame的索引,可在添加时加上
.reset_index(drop=True),例如:split_dfs.append(df.iloc[current_start:pos].copy().reset_index(drop=True))
注意事项
- 若你的连字符是全角符号
-或其他变体,需将判断条件中的'-'替换为对应字符 - 7万行数据的处理效率有保障,所有操作均基于pandas的索引切片,属于高效向量操作
内容的提问来源于stack exchange,提问作者luscgu00
相关产品推荐
相关产品推荐

