如何基于日期的小时部分向量化计算Pandas新列?
Vectorized Solution for DateTime-Based Column Calculation in Pandas
Hey there, let's ditch that slow apply() call and swap it out for proper vectorized operations—this will be way faster for your large dataset!
The problem with your attempted code is that you can’t directly access .hour or .minute on the entire Date Series. Instead, we use pandas' .dt accessor, which lets us extract datetime components (like hour, minute) in a fully vectorized way across the whole Series.
Here's the Fixed, Fast Vectorized Code
import numpy as np import pandas as pd from datetime import datetime, timedelta # Your sample DataFrame df = pd.DataFrame({'Date': np.arange(datetime(2000,1,1), datetime(2000,1,2), timedelta(hours=3)).astype(datetime)}) # Vectorized calculation of SinHour df['SinHour'] = np.sin((df['Date'].dt.hour + df['Date'].dt.minute / 60.0) * np.pi / 12.0) print(df)
Output (Matches Your Original, But Runs Way Faster)
Date SinHour 0 2000-01-01 00:00:00 0.000000e+00 1 2000-01-01 03:00:00 7.071068e-01 2 2000-01-01 06:00:00 1.000000e+00 3 2000-01-01 09:00:00 7.071068e-01 4 2000-01-01 12:00:00 1.224647e-16 5 2000-01-01 15:00:00 -7.071068e-01 6 2000-01-01 18:00:00 -1.000000e+00 7 2000-01-01 21:00:00 -7.071068e-01
Why This Works (And Is So Much Faster)
- The
.dtaccessor (df['Date'].dt.hour,df['Date'].dt.minute) returns pandas Series with datetime components, computed via optimized vectorized operations under the hood—no slow Python-level row-by-row loops. - Numpy's
np.sin()operates directly on the entire array/Series in one pass, avoiding the heavy overhead ofapply()which runs a lambda function for every single row.
For large datasets (like millions of rows), this approach will be orders of magnitude faster than your original method. You’ll see a massive difference in runtime once you scale it up.
内容的提问来源于stack exchange,提问作者Dennis Soemers
相关产品推荐
相关产品推荐

