Pandas中是否支持向量化的string.format?替代apply优化大数据集效率
apply for String Formatting Great question! You’re absolutely right that apply becomes painfully slow with large datasets—since it’s just iterating row-by-row in Python, it misses out on Pandas’ optimized, C-backed vectorized operations. Your use case is ideal for vectorization because each timestamp calculation doesn’t depend on other rows, so we can rewrite this to run way faster.
Let’s break down your original code first to replicate the logic exactly:
joined["timestamp"] = joined.apply(lambda row: args.date + " {:0>2d}:{:0>2d}:00".format(row["tid"]/6, row["tid"]%6*10), axis=1)
This takes the tid value, calculates hours as tid/6 (rounded down to an integer), minutes as tid%6 *10, formats both to 2-digit strings with leading zeros, then appends them to args.date.
The Vectorized Solution
Instead of looping through each row, we’ll use Pandas’ built-in vectorized string and arithmetic operations:
# Calculate hour and minute values using vectorized arithmetic hours = (joined['tid'] // 6).astype(int) # Integer division to get whole hours minutes = (joined['tid'] % 6 * 10).astype(int) # Format to 2-digit strings with leading zeros (vectorized str method) hour_str = hours.astype(str).str.zfill(2) minute_str = minutes.astype(str).str.zfill(2) # Combine all parts into the final timestamp joined['timestamp'] = args.date + ' ' + hour_str + ':' + minute_str + ':00'
Or if you prefer a concise one-liner (no intermediate columns needed):
joined['timestamp'] = ( args.date + ' ' + (joined['tid'] // 6).astype(str).str.zfill(2) + ':' + (joined['tid'] % 6 * 10).astype(str).str.zfill(2) + ':00' )
Why This Works (and Is Faster)
- All operations here are vectorized: Pandas handles calculations in optimized C code instead of Python-level loops, which can lead to 10-100x speedups on large datasets.
str.zfill(2)is a vectorized string method that applies leading-zero formatting to every element in the series at once—no lambda or row iteration required.- We use integer division (
//) instead of regular division (/) to avoid floating-point values, which perfectly matches the{:0>2d}format specifier’s requirement for integers.
Quick Validation
Let’s test with sample values to ensure it matches your original logic:
- If
tid = 5:5//6 = 0,5%6*10 = 50→ timestamp becomesargs.date + " 00:50:00" - If
tid = 13:13//6 = 2,13%6*10 = 10→ timestamp becomesargs.date + " 02:10:00"
This produces exactly the same output as your original apply code, but without the performance bottleneck.
内容的提问来源于stack exchange,提问作者DarkZero

