You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于行数与id条件用dplyr自动拆分数据框并导出CSV

按规则拆分大型数据集为固定行数的CSV文件

我有一个15万+行、15列的数据集,需要按以下规则自动拆分并导出为小型CSV文件:

  • 每个文件固定包含4行数据(最后一个文件行数可少于4)
  • 同一id对应的行需尽量集中在同一文件(已知每个id对应的行数不超过4)
  • 拆分后的子数据集需保留原数据的顺序与日期索引

示例数据

id=c(1,1,2,1,2,3,4,5,6,2,3)
letter=c("a","b","c","d","e","f","g","h","i","j","k")
df <- data.frame(id,letter)

按id排序后的数据:

id letter
1   1      a
2   1      b
3   1      d
4   2      c
5   2      e
6   2      j
7   3      f
8   3      k
9   4      g
10  5      h
11  6      i

期望输出示例

df_20230704_1

id letter
1  1      a
2  1      b
3  1      d
4  4      g

df_20230704_2

id letter
1  2      c
2  2      e
3  2      j
4  5      h

df_20230704_3

id letter
1  3      f
2  3      k
3  6      i

解决方案(R语言实现)

核心思路是先统计每个id的行数,再按顺序将id分组打包到容量为4的批次中,确保同一id的行不跨批次,最后按批次导出CSV。

library(tidyverse)

# 替换为你的实际数据集加载代码
# df <- read.csv("your_large_dataset.csv")

# 若需要先按id排序,取消下面注释
# df <- df %>% arrange(id)

# 统计每个id的行数
id_row_counts <- df %>% count(id, name = "row_count")

# 为每个id分配批次
current_batch <- 1
remaining_capacity <- 4
batch_mapping <- tibble(id = integer(), batch = integer())

for (i in seq_len(nrow(id_row_counts))) {
  current_id <- id_row_counts$id[i]
  id_rows <- id_row_counts$row_count[i]
  
  if (remaining_capacity >= id_rows) {
    batch_mapping <- bind_rows(batch_mapping, tibble(id = current_id, batch = current_batch))
    remaining_capacity <- remaining_capacity - id_rows
  } else {
    current_batch <- current_batch + 1
    batch_mapping <- bind_rows(batch_mapping, tibble(id = current_id, batch = current_batch))
    remaining_capacity <- 4 - id_rows
  }
}

# 将批次关联到原数据集
df_with_batch <- df %>% left_join(batch_mapping, by = "id")

# 按批次拆分并导出CSV
df_with_batch %>%
  group_by(batch) %>%
  group_walk(function(sub_df, key) {
    file_name <- sprintf("df_%s_%d.csv", format(Sys.Date(), "%Y%m%d"), key$batch)
    # 移除批次列后导出,保留原行索引
    write.csv(sub_df %>% select(-batch), file_name, row.names = TRUE, na = "")
    cat(sprintf("已导出文件:%s\n", file_name))
  })

代码说明

  1. id行数统计:提前计算每个id的行数,避免后续分配批次时出现容量溢出。
  2. 批次分配逻辑:按id的出现顺序(或排序后的顺序)打包,当前批次容量不足时自动新建批次,保证同一id的行集中在同一文件。
  3. 导出处理:生成带日期和批次号的文件名,导出时移除批次列,保留原数据的行索引和顺序。

内容的提问来源于stack exchange,提问作者Alex

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.17 04:32:41