如何在R中计算datetime变量的密度?转换后密度估计值偏低求解
Great question—let's walk through why your current approach is causing issues and how to fix it:
1. Is using as.numeric() directly on timezone-aware datetime strings appropriate?
No, and here's why:
- Timezone shifts create discontinuous values: Your data mixes EST (UTC-5) and EDT (UTC-4). When you convert these strings to numeric directly, the timezone switch will create sudden jumps in the timestamp values (e.g., a 3600-second drop when switching back from EDT to EST), which breaks the continuity assumption density estimation relies on.
- Oversized numeric range:
as.numeric(POSIXct)returns seconds since the 1970 Unix epoch. Your dates span Jan-Nov 2016, so the numeric values fall between ~1.45e9 and ~1.47e9. Since probability density integrates to 1 over the entire range, a huge range like this will naturally compress the density values to very small numbers—this is why your estimates look "low".
On top of that, you only have 10 data points, which makes density estimation inherently noisy and unreliable with default settings.
2. Step-by-step fixes
Step 1: Properly parse timezone-aware datetimes
First, convert your strings to a proper POSIXct object that handles EST/EDT automatically (they're part of the America/New_York timezone):
tenevents <- data.frame( datetime = c("2016-01-18 07:00:00 EST", "2016-02-15 11:00:00 EST", "2016-02-25 12:00:00 EST", "2016-03-01 11:00:00 EST", "2016-03-12 09:00:00 EST", "2016-05-19 16:46:00 EDT", "2016-09-13 09:30:00 EDT", "2016-09-15 07:15:00 EDT", "2016-10-03 10:26:00 EDT", "2016-11-19 04:00:00 EST") ) # Convert to POSIXct with correct timezone tenevents$datetime <- as.POSIXct(tenevents$datetime, tz = "America/New_York")
Step 2: Convert to relative time to shrink the numeric range
Instead of using epoch seconds, convert your timestamps to a relative scale (like hours since the first event, or days since the start of the year). This reduces the numeric range and makes density values more intuitive:
# Calculate hours relative to the first event tenevents$hours_since_first <- as.numeric(difftime(tenevents$datetime, min(tenevents$datetime), units = "hours")) # Or days since Jan 1, 2016 tenevents$days_since_newyear <- as.numeric(difftime(tenevents$datetime, as.POSIXct("2016-01-01", tz = "America/New_York"), units = "days"))
Step 3: Optimize density estimation for small samples
With only 10 data points, default bandwidth settings might not work well. Try these tweaks:
- Use a bandwidth tailored for small samples (like the
"sj"method) instead ofucv: - Add a rug plot to your density plot to visualize where your actual data points lie, which helps contextualize the density curve:
# Density estimate using relative hours v <- density(tenevents$hours_since_first, bw = "sj") # Plot with rug overlay plot(v, main = "Density of Event Times (Hours Since First Event)") rug(tenevents$hours_since_first) # Shows exact data positions
Bonus: If you care about time-of-day distribution
If you're interested in when events happen during the day (not across months), extract just the time component (hours with decimal minutes) and estimate density on that:
# Extract hour of day (e.g., 7:15 AM = 7.25) tenevents$hour_of_day <- as.numeric(format(tenevents$datetime, "%H")) + as.numeric(format(tenevents$datetime, "%M"))/60 # Density estimate for time-of-day, with a sensible bandwidth (1 hour) v_timeofday <- density(tenevents$hour_of_day, bw = 1) plot(v_timeofday, main = "Density of Event Times (Hour of Day)", xlim = c(0, 24)) rug(tenevents$hour_of_day)
内容的提问来源于stack exchange,提问作者dank

