如何缩减UCI异质性活动识别数据集中时间特征的数值大小?
Great question—those massive Arrival_Time and Creation_Time values can be a headache for memory usage and neural network training. Let’s walk through the best approaches to shrink their size while keeping (or even improving) your model’s performance:
1. Convert to Relative Time (Most Recommended)
Absolute timestamps like these rarely add value to activity recognition tasks—what matters is the relative timing between data points, not the exact Unix time they were recorded. Here’s how to fix this:
- Subtract the first timestamp in each sequence: For each user/activity sequence, take the first
Arrival_Timevalue and subtract it from every subsequentArrival_Timein that sequence. This gives you a relative offset starting at 0, turning values like1424696633909into small numbers like0,9,10, etc. Do the same forCreation_Time. - Calculate time differences: You could also compute the difference between consecutive timestamps (e.g.,
Arrival_Time[i] - Arrival_Time[i-1]) to capture the sampling interval, or the gap betweenArrival_TimeandCreation_Time(sensor latency). These metrics are often more useful for your model than raw absolute times.
This approach drastically reduces the numerical scale while preserving meaningful temporal information, and it’s far better than just scaling raw timestamps.
2. Normalization/Standardization (If You Must Keep Absolute Timestamps)
If for some reason you need to retain absolute timestamps, normalization is a valid way to shrink their size. But keep these rules in mind:
- Normalize all numerical features together: Yes, consistency matters here. Neural networks rely on gradient-based optimization, and features with wildly different scales (like your
-5.95to8.20acceleration values vs.1e12timestamps) will cause unstable gradient updates, slowing down training or preventing convergence. - Choose the right scaling method:
- Min-Max Normalization: Scales values to the
[0, 1]range using(value - min) / (max - min). Perfect for compressing large ranges and works well with uniform data distributions. - Z-Score Standardization: Converts values to have a mean of 0 and standard deviation of 1 using
(value - mean) / std. Better if your data has outliers, as it’s more robust to extreme values.
- Min-Max Normalization: Scales values to the
- Use training-only stats: Always compute the min/max/mean/std using only your training dataset, then apply those same values to validation and test data to avoid data leakage.
3. Optimize Data Types for Memory Savings
Even after converting to relative times or normalizing, you can cut memory usage further by adjusting the data type of these columns:
- Raw timestamps are often stored as
int64, but relative times or normalized values can usually fit into smaller types likeint32(if values are under ~2 billion) or evenfloat32(for normalized data). This halves or quarters the memory footprint of those columns without losing precision.
Should You Drop the Timestamps Entirely?
Only if you’re sure they add no value. For example, if your model is a simple MLP that treats each row as an independent sample (ignoring sequence order), you might not need them. But for sequence models like LSTMs or Transformers, temporal context (even relative timing) can improve performance—so think twice before discarding them.
内容的提问来源于stack exchange,提问作者Kristofer

