如何在R Studio中计算邮政编码的建筑类噪音投诉占比标准差
解决方法:分步计算邮编级建筑噪音投诉占比及标准差
步骤1:数据预处理与筛选
先从原始数据中筛选出所有噪音类投诉(投诉类型包含Noise:),同时标记每条投诉是否涉及建筑施工(描述中包含Construct,忽略大小写避免遗漏)。
用R代码实现:
library(dplyr) # 假设你的数据框名为nyc_noise_complaints processed_data <- nyc_noise_complaints %>% # 筛选噪音投诉 filter(grepl("Noise:", complaint_type)) %>% # 标记是否为建筑相关噪音投诉 mutate(is_construction = grepl("Construct", description, ignore.case = TRUE))
步骤2:按邮编计算建筑噪音投诉占比
按邮政编码分组,计算每组的总噪音投诉数、建筑类噪音投诉数,再算出占比:
zipcode_ratios <- processed_data %>% group_by(zipcode) %>% summarise( total_noise = n(), # 该邮编总噪音投诉数 construction_noise = sum(is_construction, na.rm = TRUE), # 建筑类噪音投诉数 construction_ratio = construction_noise / total_noise # 计算占比 )
步骤3:计算所有邮编占比的标准差
zipcode_ratios中已包含每个邮编的建筑噪音投诉占比construction_ratio,直接用sd()函数计算标准差即可:
ratio_sd <- sd(zipcode_ratios$construction_ratio, na.rm = TRUE) # 打印结果 cat("所有邮编建筑噪音投诉占比的标准差为:", ratio_sd, "\n")
针对你之前尝试的补充说明
你之前按邮编+描述分组统计数量的方式,会把每个不同描述拆分开,不利于直接计算占比。我们需要的是每个邮编的整体统计:先统计该邮编下所有噪音投诉的总数,再统计其中符合建筑描述的数量,两者相除得到占比,最后对所有占比求标准差即可。
示例数据验证
用模拟的纽约噪音投诉样本测试:
# 模拟样本数据 sample_data <- tibble( zipcode = c(10001, 10001, 10001, 10002, 10002, 10003, 10003, 10003, 10003), complaint_type = c(rep("Noise: Construction", 3), rep("Noise: Street", 2), rep("Noise: Construction", 2), rep("Noise: Street", 2)), description = c("Loud construction at night", "Construction noise early morning", "Jackhammer noise", "Car honking", "Bar music", "Weekend construction", "Construction after hours", "Sirens", "Traffic noise") ) # 执行上述步骤 processed_sample <- sample_data %>% filter(grepl("Noise:", complaint_type)) %>% mutate(is_construction = grepl("Construct", description, ignore.case = TRUE)) zip_ratios_sample <- processed_sample %>% group_by(zipcode) %>% summarise( total_noise = n(), construction_noise = sum(is_construction), construction_ratio = construction_noise / total_noise ) sd(zip_ratios_sample$construction_ratio) # 输出结果约为0.4714
内容的提问来源于stack exchange,提问作者Data Science Analytics Manager
相关产品推荐
相关产品推荐

