如何按国家分组计算1955年与1952年IV值的比值?
高效计算各国IV值年份比值并生成新数据集
原始数据集
data <- structure(list(country = c("Poland", "Poland", "Poland", "Poland", "Poland", "Poland", "Portugal", "Portugal", "Portugal", "Portugal", "Portugal", "Portugal", "Spain", "Spain", "Spain", "Spain", "Spain", "Spain"), Code = c("POL", "POL", "POL", "POL", "POL", "POL", "PRT", "PRT", "PRT", "PRT", "PRT", "PRT", "ESP", "ESP", "ESP", "ESP", "ESP", "ESP"), year = c(1950, 1951, 1952, 1953, 1954, 1955, 1950, 1951, 1952, 1953, 1954, 1955, 1950, 1951, 1952, 1953, 1954, 1955), IV = c(3, 3, 3, 3, 3, 3, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 2)), row.names = c(1L, 2L, 3L, 4L, 5L, 6L, 7L, 8L, 9L, 10L, 11L, 12L, 13L, 14L, 15L, 16L, 17L, 18L), class = "data.frame")
需求说明
生成新数据集newdata,仅保留country列和新变量new.variable,其中new.variable为每个国家1955年的IV值除以1952年的IV值,删除原数据集其他行和列,要求高效实现(适配大量国家分组的场景)。
高效实现方法
方法1:使用dplyr(tidyverse生态)
适合习惯tidy语法的用户,代码可读性强,处理中等规模数据效率足够:
library(dplyr) newdata <- data %>% group_by(country) %>% summarise( new.variable = IV[year == 1955] / IV[year == 1952] ) %>% ungroup()
方法2:使用data.table(高性能大数据处理)
适合处理超大规模数据集,运算速度更快:
library(data.table) # 转换为data.table格式 setDT(data) newdata <- data[, .(new.variable = IV[year == 1955] / IV[year == 1952]), by = country]
结果验证
运行上述代码后,newdata的输出结果如下:
print(newdata) # country new.variable # 1 Poland 1.0 # 2 Portugal 1.0 # 3 Spain 2.0
内容的提问来源于stack exchange,提问作者Rustam
相关产品推荐
相关产品推荐

