You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在出现连字符时拆分DataFrame并删除对应行

拆分含连字符的DataFrame并删除对应行

核心思路

先定位所有包含连字符-的行,再以这些行为分隔点,将原DataFrame切割成多个连续的、不含连字符的子DataFrame,同时丢弃含连字符的行。

代码实现

1. 读取数据

假设你的数据是分号分隔格式,先读取为pandas DataFrame:

import pandas as pd

# 替换为你的数据路径或读取逻辑
df = pd.read_csv('your_data_file.csv', sep=';', header=0)

2. 定位含连字符的行

遍历每行,标记并提取包含连字符的行索引:

# 把每行转为字符串后检查是否存在连字符'-'
has_hyphen = df.apply(lambda row: '-' in row.astype(str).values, axis=1)
# 获取所有含连字符行的索引列表
hyphen_positions = df[has_hyphen].index.tolist()

3. 拆分DataFrame

根据定位到的索引切割原DataFrame,生成子DataFrame列表:

split_dfs = []
current_start = 0

for pos in hyphen_positions:
    # 提取从current_start到pos前一行的有效数据
    if current_start < pos:
        split_dfs.append(df.iloc[current_start:pos].copy())
    # 更新起始位置为当前含连字符行的下一行
    current_start = pos + 1

# 处理最后一段未被切割的有效数据
if current_start < len(df):
    split_dfs.append(df.iloc[current_start:].copy())

结果说明

  • 执行后,split_dfs就是拆分后的所有子DataFrame列表,对应你示例中的预期结果(2个DataFrame)
  • 如果需要重置每个子DataFrame的索引,可在添加时加上.reset_index(drop=True),例如:
    split_dfs.append(df.iloc[current_start:pos].copy().reset_index(drop=True))
    

注意事项

  • 若你的连字符是全角符号-或其他变体,需将判断条件中的'-'替换为对应字符
  • 7万行数据的处理效率有保障,所有操作均基于pandas的索引切片,属于高效向量操作

内容的提问来源于stack exchange,提问作者luscgu00

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.15 18:40:32