如何在Pandas中将数字格式时间转换为时段并新增列?
Absolutely, your approach is totally feasible and actually a straightforward, efficient way to handle this time-to-period conversion! Extracting the first two characters (which represent the hour) and using conditional logic is exactly the right direction here—let's break down how to implement it cleanly in pandas.
Step 1: Ensure your time column is string type
First, make sure your time values are treated as strings (since you're slicing characters). If they're stored as integers, convert them first to preserve leading zeros (like 0325 instead of 325):
df['time_column'] = df['time_column'].astype(str)
If some values are missing leading zeros (e.g., 300 instead of 0300), pad them with str.zfill(4):
df['time_column'] = df['time_column'].str.zfill(4)
Step 2: Extract the hour component
Slice the first two characters of each string, then convert to an integer for numerical comparison:
df['hour'] = df['time_column'].str[:2].astype(int)
Step 3: Map hours to periods
You have two clean options to assign period labels based on the extracted hour:
Option 1: Chained np.where for simple conditions
Perfect if you have a small number of intervals and want explicit control:
import numpy as np # Define your period rules (adjust thresholds to match your needs) df['period'] = np.where( df['hour'] < 6, 'Night', np.where(df['hour'] < 12, 'Morning', np.where(df['hour'] < 18, 'Afternoon', 'Evening')) )
Option 2: pd.cut for cleaner interval handling
Ideal if you prefer a more readable way to define bins:
bins = [-1, 6, 12, 18, 24] labels = ['Night', 'Morning', 'Afternoon', 'Evening'] df['period'] = pd.cut(df['hour'], bins=bins, labels=labels, include_lowest=True)
include_lowest=True ensures values like 0 (midnight) are included in the first bin.
Example Output
For a sample DataFrame, you’ll get results like this:
| time_column | hour | period |
|---|---|---|
| 1300 | 13 | Afternoon |
| 0325 | 3 | Night |
| 1005 | 10 | Morning |
| 1830 | 18 | Evening |
Key Notes
- Adjust the hour thresholds to match your exact definition of each period (there’s no universal standard—tweak based on your use case).
- This method is efficient even for large DataFrames, since pandas string operations and vectorized conditionals are optimized for performance.
Your initial idea is solid—slicing the hour component and using conditional logic is a pragmatic, easy-to-maintain solution for this task.
内容的提问来源于stack exchange,提问作者user20977195

