使用as.POSIXct转换时列值变为NA,求R语言解决方案
as.POSIXct() Returning All NAs in Real Dataset Let's break down why your time parsing line is failing with real data (but working for simulated data) and fix it step by step—this almost always comes down to mismatched time formats between your test and actual files.
Step 1: Diagnose the Root Cause
First, let's inspect what your real data's X2 column actually looks like. Add these lines right after reading in your CSV:
dat <- read.csv(fn, stringsAsFactors=FALSE) head(dat$X2) # Check the first few time values table(nchar(dat$X2)) # See if all time strings are the same length (e.g., 5 chars for HH:MM) print(thisdate) # Confirm your date extraction is working correctly
Common issues here include:
- Your real
X2includes seconds (e.g.,00:00:00instead of00:00), so your%H:%Mformat doesn't match. - Some entries use 12-hour time with AM/PM (e.g.,
10:05 AMinstead of10:05). - Leading zeros are missing (e.g.,
1:05instead of01:05—whileas.POSIXctusually handles this, it's worth checking). - The
thisdateextraction is broken (e.g., pulling an invalid string instead ofYYYYMMDD).
Step 2: Fix the Time Parsing
Option 1: Adjust the format Parameter to Match Real Data
If you find your X2 has seconds, update the format string to include %S:
dat$X2 <- as.POSIXct(paste(thisdate, dat$X2), format = "%Y%m%d %H:%M:%S")
If it uses 12-hour time with AM/PM, use this format instead:
dat$X2 <- as.POSIXct(paste(thisdate, dat$X2), format = "%Y%m%d %I:%M %p")
Option 2: Use lubridate for Flexible Automatic Parsing
For more robustness (especially if your data has inconsistent formats), use the lubridate package—it handles most common time patterns without needing to specify exact formats:
# Install first if you haven't: install.packages("lubridate") library(lubridate) # For HH:MM times dat$X2 <- ymd_hm(paste(thisdate, dat$X2)) # If seconds are present: dat$X2 <- ymd_hms(paste(thisdate, dat$X2))
If there are multiple possible formats in your data, use parse_date_time to try several patterns:
dat$X2 <- parse_date_time(paste(thisdate, dat$X2), orders = c("Ymd HM", "Ymd HMS", "Ymd IMS p"))
Step 3: Verify and Clean Up
After parsing, check for remaining NAs and identify why they failed:
na_rows <- which(is.na(dat$X2)) print(dat$X2[na_rows]) # Inspect the problematic time strings
This will help you catch edge cases like typos in time values that need manual cleaning.
Why Your Previous Workarounds Failed
- Using
seq(twotimes[1], twotimes[2], by="min")directly assigns 1440 rows todat$X2, but your original data has fewer rows—hence the "number of rows doesn't match" error. - Merging with a raw time vector (
merge(dat, time)) creates a Cartesian product (every row indatpaired with every time), which is why you got ~2 million rows. Your originalmergewithallminuteswas correct—you just needed to fix the time parsing first.
内容的提问来源于stack exchange,提问作者MT32

