如何为大型DataFrame快速将时间值映射为对应班次(Shift)?
If you're working with a large DataFrame, vectorized operations are key to keeping things fast—avoid loops at all costs! Here's how to convert your time values to the corresponding shift labels efficiently:
Step 1: Ensure Your Time Column is in a Datetime Format
First, convert your string-based Time column to a pandas datetime object (this makes extracting hours straightforward):
import pandas as pd import numpy as np # Sample DataFrame matching your structure data = { 'Row': [1,2,3,4,5,6,7], 'Time': ['01:15:12','09:18:22','21:56:01','13:33:23','11:59:56','17:08:38','21:55:16'] } df = pd.DataFrame(data) # Convert Time column to datetime (ignores date, focuses on time) df['Time'] = pd.to_datetime(df['Time'], format='%H:%M:%S')
Step 2: Map Hours to Shifts Using Vectorized Logic
We have two great, efficient options here—both avoid slow row-wise operations:
Option 1: Using pd.cut (Clean and Concise)
This method bins hour values directly into your predefined intervals:
# Define bins (0-8, 8-16, 16-24) and corresponding shift labels bins = [0, 8, 16, 24] labels = [1, 2, 3] # Extract hour from datetime and apply binning df['Shift'] = pd.cut(df['Time'].dt.hour, bins=bins, labels=labels, include_lowest=True)
include_lowest=Trueensures00:00:00is correctly grouped into Shift 1.- Bins are exclusive of upper bounds, so
08:00:00falls into Shift 2 and16:00:00into Shift 3—perfect for your rules.
Option 2: Using np.select (Flexible for Complex Rules)
If you need explicit control over conditions (great if rules ever change), use np.select:
# Define conditions for each shift conditions = [ df['Time'].dt.hour < 8, (df['Time'].dt.hour >= 8) & (df['Time'].dt.hour < 16), df['Time'].dt.hour >= 16 ] # Corresponding shift values choices = [1, 2, 3] # Apply conditions to create Shift column df['Shift'] = np.select(conditions, choices)
Step 3: Clean Up (Optional)
If you don't want the default date (1900-01-01) attached to your Time values, convert it back to a time string:
df['Time'] = df['Time'].dt.time
Final Result
Your DataFrame will now look exactly like what you requested:
Row Time Shift 0 1 01:15:12 1 1 2 09:18:22 2 2 3 21:56:01 3 3 4 13:33:23 2 4 5 11:59:56 2 5 6 17:08:38 3 6 7 21:55:16 3
Both methods leverage pandas/numpy's optimized vectorized operations, making them ideal for large datasets—they’ll run orders of magnitude faster than manual row loops.
内容的提问来源于stack exchange,提问作者Hussain Rahiminejad

