You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Pandas中是否支持向量化的string.format?替代apply优化大数据集效率

Vectorized Alternative to Pandas apply for String Formatting

Great question! You’re absolutely right that apply becomes painfully slow with large datasets—since it’s just iterating row-by-row in Python, it misses out on Pandas’ optimized, C-backed vectorized operations. Your use case is ideal for vectorization because each timestamp calculation doesn’t depend on other rows, so we can rewrite this to run way faster.

Let’s break down your original code first to replicate the logic exactly:

joined["timestamp"] = joined.apply(lambda row: args.date + " {:0>2d}:{:0>2d}:00".format(row["tid"]/6, row["tid"]%6*10), axis=1)

This takes the tid value, calculates hours as tid/6 (rounded down to an integer), minutes as tid%6 *10, formats both to 2-digit strings with leading zeros, then appends them to args.date.

The Vectorized Solution

Instead of looping through each row, we’ll use Pandas’ built-in vectorized string and arithmetic operations:

# Calculate hour and minute values using vectorized arithmetic
hours = (joined['tid'] // 6).astype(int)  # Integer division to get whole hours
minutes = (joined['tid'] % 6 * 10).astype(int)

# Format to 2-digit strings with leading zeros (vectorized str method)
hour_str = hours.astype(str).str.zfill(2)
minute_str = minutes.astype(str).str.zfill(2)

# Combine all parts into the final timestamp
joined['timestamp'] = args.date + ' ' + hour_str + ':' + minute_str + ':00'

Or if you prefer a concise one-liner (no intermediate columns needed):

joined['timestamp'] = (
    args.date + ' '
    + (joined['tid'] // 6).astype(str).str.zfill(2)
    + ':'
    + (joined['tid'] % 6 * 10).astype(str).str.zfill(2)
    + ':00'
)

Why This Works (and Is Faster)

  • All operations here are vectorized: Pandas handles calculations in optimized C code instead of Python-level loops, which can lead to 10-100x speedups on large datasets.
  • str.zfill(2) is a vectorized string method that applies leading-zero formatting to every element in the series at once—no lambda or row iteration required.
  • We use integer division (//) instead of regular division (/) to avoid floating-point values, which perfectly matches the {:0>2d} format specifier’s requirement for integers.

Quick Validation

Let’s test with sample values to ensure it matches your original logic:

  • If tid = 5: 5//6 = 0, 5%6*10 = 50 → timestamp becomes args.date + " 00:50:00"
  • If tid = 13: 13//6 = 2, 13%6*10 = 10 → timestamp becomes args.date + " 02:10:00"

This produces exactly the same output as your original apply code, but without the performance bottleneck.

内容的提问来源于stack exchange,提问作者DarkZero

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 06:54:45