按资源分组计算最常用单位下的总数量(R语言数据框处理)
按资源分组计算最常用单位下的总数量
先给出原始数据:
df <- data.frame(resource = c("gold", "gold", "gold", "silver", "silver", "gold", "silver", "bronze"), amount = c(500, 2000, 4, 8, 100, 2000, 3, 5), unit = c("g", "g", "kg", "ton", "kg", "g", "ton", "kg"), price = c(10, 10, 10000, 50000, 50, 10, 50000, 20))
这里用dplyr包实现需求,完全贴合你提出的三步思路:
library(dplyr) final_result <- df %>% # 1. 计算单条记录价值,再按资源分组求和得到总价值 mutate(item_value = amount * price) %>% group_by(resource) %>% summarise( total_value = sum(item_value), # 2. 统计各单位出现频率,取最高频的(频率相同时随机选第一个) top_unit = names(sort(table(unit), decreasing = TRUE))[1], # 获取该单位对应的单价(假设同一资源同一单位的单价一致) unit_price = price[unit == top_unit][1] ) %>% # 3. 用总价值除以对应单价得到转换后的数量,调整字段名匹配预期结果 mutate( amount = total_value / unit_price, value = total_value ) %>% select(resource, value, price = unit_price, unit = top_unit, amount) %>% ungroup() print(final_result)
运行后输出:
# A tibble: 3 × 5 resource value price unit amount <chr> <dbl> <dbl> <chr> <dbl> 1 bronze 100 20 kg 5 2 gold 85000 10 g 8500 3 silver 555000 50000 ton 11.1
注:原始数据中bronze仅有1条记录(5kg,单价20),总价值为100,按单价20转换后数量应为5,推测你给出的预期结果中bronze的amount是笔误。
内容的提问来源于stack exchange,提问作者Anton
相关产品推荐
相关产品推荐

