R脚本分段运行正常但整段运行报`-.POSIXt`错误求助
问题描述
整段运行R脚本时立即终止并返回如下错误:
Error in `-.POSIXt`(`*tmp*`) : unary '-' is not defined for "POSIXt" objects
但分段运行(例如先执行Q4部分,再执行Q5部分等)时完全正常。相关脚本如下:
# Installing packages I will need install.packages("dplyr") install.packages("tidyr") install.packages("readxl") install.packages("psych") install.packages("car") # loading packages library(dplyr) library(tidyr) library(readxl) library(psych) library(car) # Importing the Data Frame aqdata <- read_xlsx("D:\\tVNS - Data Folder\\AnalitiQs\\Company_Y_base_data.xlsx") # Inspecting the head, colnames and nrows to see that everything was imported successfully head(aqdata) colnames(aqdata) nrow(aqdata) # Using distinct to compare nrows to ensure that all rows are unique distinct(aqdata, Employee_ID) -------------- # Q4 Step 1: # Converting the stop Month column to a date vector aqdata$Stop_Month <- as.Date(aqdata$Stop_Month) # Q4 Step 2 - Filtering employees by department & missing values (Q depends on if last working day is date of Stop_Month or (as I interpret) last day of work was 1 day before Stop_Month date) workers_post28 <- aqdata %>% filter(Department == "Operations" & is.na(Stop_Month)) # Q4 Step 3: Final Answer num_rows4 <- nrow(workers_post28) message <- paste("On the 28th of July 2020", num_rows4, "people worked within the Operations Department of company Y") print(message) -------------- # Q5 Step 1: Converting the Start Month column to a date vector aqdata$Start_Month <- as.Date(aqdata$Start_Month) # Q5 Step 2: Filtering employees by dep, permanent contract. & 28th of Dec 2019 workers_on28Dec2019 <- aqdata %>% filter(Department == "Operations" & Type_of_Contract == "Permanent") %>% # Q5 Step 3: Filtering resulting employees by condition start date before Dec 28th 2019 and Stop date after 28th of Dec 2019 (based on prev question interpretation) filter(Start_Month < as.Date("2019-12-28") & is.na(Stop_Month) | Stop_Month > as.Date("2019-12-28")) # Q5 Step 4: Final Answer num_rows5 <- nrow(workers_on28Dec2019) messageQ5 <- paste("On the 28th of Dec 2019", num_rows5, "people in the Operations department had a permanent contract") print(messageQ5) # Q5 Step 3.1: Because I was unsure if my result was correct, I decided to verify that no employee not satisfying my conditions were left in the dataframe. # Filter to only include employees who started working after December 28th, 2019 before_28dec2019 <- workers_on28Dec2019 %>% filter(Start_Month > as.Date("2019-12-28")) # Count the number of rows in the final result num_rowscheck <- nrow(final_result) # Print the result if (num_rowscheck == 0) { print("The resulting dataframe does not contain any employees who started working before December 28th, 2019") } -------------- # Q6 Step 1: Calculate sum of distance and dividing by length of column (Can also use mean() function for more concise code) AverageDistance <- sum(aqdata$Distance_Work_KM) / length(aqdata$Distance_Work_KM) messageQ6 <- paste("The average distance to work for any employee who has ever worked for Company Y is ", AverageDistance, "KM") print(messageQ6) -------------- # Q7 Step 1: Filter workers by gender with conditions satisfying working on 2020-07-28 --> Creating the groups Q7df <- aqdata %>% # Step 1.1 Filter rows based on conditions of the question (using same interpretation from Q4 where NA in Stop_Month indicates that they were working on the 28th as this is their last working day) filter(Start_Month < as.Date("2020-07-28") & is.na(Stop_Month)) %>% # Step 1.2 Group by gender group_by(Gender) %>% # Step 1.3 Select Salary column to incl into subset DF select(Gender,Salary) # Q7 Step 2: Assumption checking first normality, then equality of variances describe(Q7df$Salary) # Getting M, SD and Kurtosis values for Normality assumption --> Skew/Kurtosis do not fall within -2 or +2 so its violated shapiro.test(Q7df$Salary) # Another way of testing the assumption --> Also indicates violation. # Both test indicate that the assumption is violated but our sample is 319 so it should be okay # Testing equality of variances leveneTest(Q7df$Salary ~ Q7df$Gender) # The test is non-significant meaning that the assumption is not violated # Q7 Step 3: Running the t-test Outcome <- t.test(Salary ~ Gender, data = Q7df) # Assuming equal variances b.c leven's was non-sig. # Q7 Step 4: Final conclusions print(Outcome$estimate) print(Outcome$statistic) print(Outcome$p.value) messageQ7 <- paste("The T-test indicates that there is significant differences between the means of Female/Male in terms of Salary") print(messageQ7) --------------
原因分析与解决方案
核心原因
脚本中用来分隔各部分的--------------是无效的R语法,整段运行时R会将其解析为连续的一元负号运算符(即反复对前一行代码的执行结果取负)。如果前一行的结果是POSIXt类型的日期对象,就会触发unary '-' is not defined for "POSIXt" objects错误。而分段运行时,你手动跳过了这些分隔线,因此不会触发错误。
另外还有一个隐藏问题:Q5部分的num_rowscheck <- nrow(final_result)中,final_result变量未定义,整段运行到此处也会报错。
解决步骤
- 替换无效分隔线:将所有
--------------改为注释格式,比如# ------------------------------,这样R会忽略这些行,不会尝试执行。 - 修复未定义变量:把Q5中的
num_rowscheck <- nrow(final_result)改为num_rowscheck <- nrow(before_28dec2019),因为before_28dec2019才是前面代码生成的验证用数据框。 - 统一日期转换逻辑:建议在数据导入后一次性完成日期列的转换,避免分散在各个问题模块重复操作,示例代码如下:
# 统一转换日期列 aqdata <- aqdata %>% mutate( Stop_Month = as.Date(Stop_Month), Start_Month = as.Date(Start_Month) )
额外优化建议
- 避免重复安装包:可以添加判断逻辑,只在包未安装时才执行安装操作:
# 检查并安装所需包 required_pkgs <- c("dplyr", "tidyr", "readxl", "psych", "car") for(pkg in required_pkgs) { if(!require(pkg, character.only = TRUE)) { install.packages(pkg) library(pkg, character.only = TRUE) } }
- 模拟全新运行环境:整段运行前清空工作区(RStudio中点击
Session -> Clear Workspace),可以提前发现分段运行时因环境变量残留而隐藏的错误。
内容的提问来源于stack exchange,提问作者Boxingday
相关产品推荐
相关产品推荐

