R脚本mutate函数报错及多对象未找到问题求助
R脚本错误排查与修正方案
核心错误原因
- nottem数据集处理逻辑错误:原脚本直接将时间序列(ts)对象
nottem转成tibble,丢失了时间索引信息,导致后续mutate中调用floor_date、year()等日期函数时,因无有效日期对象触发as.POSIXlt.default()错误,后续依赖nottem_tidy_df的代码全部因对象未创建而失败。 - Titanic数据集操作遗漏与变量名错误:未执行
uncount()操作展开频数数据,导致幸存者比例计算错误;最后绘图时误用未定义的class_counts对象,实际应为class_totals_df。
分步修正方案
1. 修正nottem数据集处理流程
使用tsibble::as_tsibble()将时间序列转换为带日期索引的tibble,再提取年、月信息:
# 将nottem时间序列转为带日期索引的tsibble nottem_ts <- as_tsibble(nottem) %>% rename(temperature = value) # 重命名默认的value列为temperature # 生成整洁格式的数据集 nottem_tidy_df <- nottem_ts %>% mutate(year = year(index), month = month(index)) %>% rename(date = index) %>% # 将index重命名为date select(date, year, month, temperature)
2. 修正Titanic数据集操作
- 补充
uncount()操作展开频数数据 - 修正绘图对象名:
# 展开Titanic的频数数据并转换变量类型 titanic_factors_df <- titanic_tibble_df %>% uncount(n) %>% # 执行uncount操作 mutate(across(c(Class, Age, Sex, Survived), as.factor)) # 修正绘图对象名 ggplot(class_totals_df, aes(x = Class, y = prop_survived)) + geom_bar(stat = "identity") + scale_y_continuous(limits = c(0, 1), labels = scales::percent_format()) + labs(x = "舱位等级", y = "幸存乘客比例", title = "不同舱位乘客幸存比例") + ggsave("proportion_survived_by_class.png")
完整修正脚本
# ------------------------------ # 姓名:[你的姓名] # 作业名称:数据分析作业 # ------------------------------ # 检查并安装pacman包 if (!require("pacman")) install.packages("pacman") print("Pacman 已安装/加载") # 加载所需包 pacman::p_load(pacman, datasets, tidyverse, tsibble, lubridate) print("所有依赖包加载完成") # 处理nottem气温数据集 # 1. 加载数据集并转为带日期索引的tsibble data("nottem") print("nottem数据集加载完成") nottem_ts <- as_tsibble(nottem) %>% rename(temperature = value) print("nottem已转为带日期索引的tsibble") # 2. 生成整洁格式数据集 nottem_tidy_df <- nottem_ts %>% mutate(year = year(index), month = month(index)) %>% rename(date = index) %>% select(date, year, month, temperature) print("nottem整洁数据集创建完成") # 3. 计算年度平均气温 average_temp_by_year_df <- nottem_tidy_df %>% group_by(year) %>% summarize(avg_temp = mean(temperature, na.rm = TRUE)) print("年度平均气温数据集创建完成") # 4. 绘制年度气温趋势图并保存 ggplot(average_temp_by_year_df, aes(year, avg_temp)) + geom_line(color = "#2c3e50") + geom_smooth(method = "loess", color = "#e74c3c", se = FALSE) + ggtitle("诺丁汉年度平均气温趋势") + xlab("年份") + ylab("平均气温 (°C)") + theme_minimal() ggsave("Annual_Temperature_by_Year.png", dpi = 300) print("年度气温趋势图已保存") # 处理Titanic数据集 # 1. 加载数据集并转为tibble data("Titanic") print("Titanic数据集加载完成") titanic_tibble_df <- as_tibble(Titanic) print("Titanic已转为tibble格式") # 2. 展开频数数据并转换变量类型为因子 titanic_factors_df <- titanic_tibble_df %>% uncount(n) %>% mutate(across(c(Class, Age, Sex, Survived), as.factor)) print("Titanic数据集已展开并完成变量类型转换") # 3. 计算总体幸存比例 num_survived <- sum(titanic_factors_df$Survived == "Yes") num_total <- nrow(titanic_factors_df) prop_survived <- num_survived / num_total print(paste("总体幸存比例:", round(prop_survived, 4))) # 4. 统计各舱位总人数 class_count_df <- titanic_factors_df %>% group_by(Class) %>% summarize(total_count = n()) print("各舱位总人数统计完成") # 5. 统计各舱位幸存人数 class_survived_df <- titanic_factors_df %>% filter(Survived == "Yes") %>% group_by(Class) %>% summarize(survival_count = n()) print("各舱位幸存人数统计完成") # 6. 合并总人数与幸存人数并计算幸存比例 class_totals_df <- class_count_df %>% left_join(class_survived_df, by = "Class") %>% mutate(prop_survived = survival_count / total_count) print("舱位幸存比例计算完成") # 7. 绘制各舱位幸存比例柱状图并保存 ggplot(class_totals_df, aes(x = Class, y = prop_survived)) + geom_bar(stat = "identity", fill = "#3498db") + scale_y_continuous(limits = c(0, 1), labels = scales::percent_format()) + labs(x = "舱位等级", y = "幸存乘客比例", title = "泰坦尼克号不同舱位乘客幸存比例") + theme_minimal() ggsave("proportion_survived_by_class.png", dpi = 300) print("舱位幸存比例图已保存")
内容的提问来源于stack exchange,提问作者Stephen Magnus
相关产品推荐
相关产品推荐

