You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

R语言实现:跨年度更新县域纪念碑数量统计

问题:计算县域各年份当前纪念碑数量

问题背景

现有包含以下字段的DataFrame:

  • CountyState:县域标识
  • YearCount:统计年份
  • CountPerCounty:县域初始纪念碑总数
  • RemovalsPerCountyYear:单年度移除的纪念碑数量
  • RemovalYear:移除行为发生的年份

需要生成CurrentCount字段,统计每个县域在任意年份的当前纪念碑数量。原有的单条件ifelse方法仅能处理单一年份有移除记录的情况,无法应对多年份移除的场景。

目标DataFrame示例

最终期望的输出结构如下:

structure(list(washington = c("DC, Washington County", "DC, Washington County", 
"DC, Washington County", "DC, Washington County", "DC, Washington County", 
"DC, Washington County", "DC, Washington County", "DC, Washington County"
), year = c("2015", "2016", "2017", "2018", "2019", "2020", "2021", 
"2022"), CurrentCount = c("10", "10", "10", "10", "10", "8", 
"8", "7")), class = "data.frame", row.names = c(NA, -8L))

以华盛顿特区华盛顿县为例:初始有10座纪念碑,2020年移除2座,2022年移除1座,因此2020-2021年当前数量为8,2022年为7。

尝试的无效代码

以下代码仅适用于单年份移除的场景,多移除年份时计算结果错误:

df8$CountPerCountyYear <- ifelse(df8$RemovalYear <= df8$YearCount, df8$CountPerCounty - df8$RemovalsPerCountyYear, df8$CountPerCounty)

示例输入数据

用于测试的输入DataFrame:

structure(list(CountyState = c("DC, Washington County", "DC, Washington County", 
"DC, Washington County", "DC, Washington County", "DC, Washington County", 
"DC, Washington County", "DC, Washington County", "DC, Washington County", 
"DC, Washington County", "DC, Washington County", "DC, Washington County", 
"DC, Washington County", "DC, Washington County", "DC, Washington County", 
"DC, Washington County"), YearCount = c("2020", 
"2020", "2021", "2021", "2022", "2022", "2017", "2017", "2018", 
"2018", "2019", "2019", "2015", "2015", "2016", "2016"), CountPerCounty = c(10, 
10, 10, 10, 10, 10, 10, 10, 10, 10, 10, 10, 10, 10, 10, 10), 
    RemovalsPerCountyYear = c(2, 1, 2, 1, 2, 1, 2, 1, 2, 1, 2, 
    1, 2, 1, 2, 1), RemovalYear = c("2020", "2022", "2020", "2022", 
"2020", "2022", "2020", "2022", "2020", "2022", "2020", "2022", 
"2020", "2022", "2020", "2022")), row.names = c(14L, 15L, 
16L, 17L, 18L, 19L, 179L, 180L, 181L, 182L, 183L, 184L, 324L, 
325L, 326L, 327L), class = "data.frame")

解决方案

使用dplyr分组统计累计移除数量,再计算当前纪念碑数量:

library(dplyr)

# 1. 将年份字段转换为数值型,避免字符比较错误
df8 <- df8 %>%
  mutate(YearCount = as.numeric(YearCount),
         RemovalYear = as.numeric(RemovalYear))

# 2. 分组计算当前纪念碑数量
result_df <- df8 %>%
  group_by(CountyState, YearCount) %>%
  summarize(
    # 提取初始纪念碑总数(每组内值一致,取第一个即可)
    CountPerCounty = first(CountPerCounty),
    # 计算截至当前统计年份的累计移除数量
    TotalRemovals = sum(ifelse(RemovalYear <= YearCount, RemovalsPerCountyYear, 0)),
    # 初始总数减去累计移除数得到当前数量
    CurrentCount = CountPerCounty - TotalRemovals
  ) %>%
  ungroup()

# 查看结果
print(result_df)

代码说明

  • 类型转换:将YearCount和RemovalYear转为数值型,确保年份比较的准确性
  • 分组统计:按县域和统计年份分组,对每组内的所有移除记录判断是否发生在当前年份及之前,符合条件的移除数量求和得到累计移除数
  • 计算当前数量:用初始纪念碑总数减去累计移除数,得到该年份的当前纪念碑数量

该方案可正确处理多年份移除的场景,比如示例中华盛顿县2022年的累计移除数为2+1=3,当前数量为10-3=7,与目标结果一致。

内容的提问来源于stack exchange,提问作者Max Primbs

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.13 01:52:01