You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

R语言:计算指定行占列总和占比并筛选列的问题求助

问题背景与需求
  • 处理无法用Excel打开的大型CSV文件:约22000列、210行,带行列表头
  • 核心需求:
    1. 计算行1、47-56区间行、156-158区间行的数值之和,占除前两列外各列总和的比例
    2. 筛选并删除占比超过0.1%的列
  • 使用vroom替代readr加载文件,避免内存溢出问题
示例文件结构
H e a d e r           A1    A2    A3    A4    A5    A6
H    1    sample a     1     0     0     13    0     9
e    2    sample b     4     0     0     8     312   24
a    3    sample c     0     20    0     49    0     17
d    4    sample d     2     0     213   18    56    3
e    5    sample e     5     4     0     10    94    62
r    6    sample f     9     87    0     2     33    90
原代码与报错信息

原尝试代码:

library(dplyr)
library(vroom)

myData <- vroom("File.csv")

myData$newRow <- 100*(colSums(myData[-1, -2])/rowSums(myData[1, 47:56, 156:158]))

报错信息:

> myData$newRow <- 100*(colSums(myData[-1, -2])/rowSums(myData[1, 47:56, 156:158]))
Error:
! Assigned data `100 * ...` must be compatible with existing data.
✖ Existing data has 200 rows.
✖ Assigned data has 21941 rows.
ℹ Only vectors of size 1 are recycled.
Run `rlang::last_trace()` to see where the error occurred.
Warning message:
In drop && length(xo) == 1L :
  'length(x) = 3 > 1' in coercion to 'logical(1)'
错误原因
  1. 索引逻辑错误:myData[1, 47:56, 156:158]是三维索引,不符合data.frame的二维索引规则,正确指定多行应该用myData[c(1, 47:56, 156:158), ]
  2. 维度不匹配:colSums(myData[-1, -2])得到的是各列总和(长度等于列数-2),错误的rowSums调用返回异常维度结果,最终比例向量长度与原数据行数不匹配,导致赋值失败
  3. 需求偏差:原代码试图新增列存储比例,但实际需求是计算列层面比例并筛选列,无需修改原数据的行/列结构
修正后的解决方案代码
library(dplyr)
library(vroom)

# 1. 加载大型CSV文件
myData <- vroom("File.csv")

# 2. 定义需要计算总和的目标行索引
target_rows <- c(1, 47:56, 156:158)

# 3. 提取数值列(排除前两列表头列),分别计算目标行的列总和、所有行的列总和
numeric_cols <- myData[, -(1:2)]
target_col_sum <- colSums(numeric_cols[target_rows, ])
total_col_sum <- colSums(numeric_cols)

# 4. 计算占比(添加极小值避免除以0错误)
proportion <- 100 * target_col_sum / (total_col_sum + 1e-10)

# 5. 筛选保留占比<=0.1%的列,拼接前两列得到最终结果
filtered_data <- myData[, c(1, 2, which(proportion <= 0.1))]

# 6. 高效保存处理后的文件(可选)
vroom_write(filtered_data, "Filtered_File.csv")
代码说明
  • target_rows:统一指定目标行索引,避免多维索引错误
  • numeric_cols:单独提取数值列,聚焦核心计算逻辑
  • 除法时添加1e-10:防止某列全为0导致的除以0报错
  • which(proportion <= 0.1):精准筛选符合保留条件的列索引,结合前两列生成最终数据集
  • vroom_write:保持与vroom加载一致的高效IO,适配大型文件

内容的提问来源于stack exchange,提问作者Jay

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.25 06:13:31