You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Pandas循环优化需求:日志DataFrame特定记录提取的循环改进

Optimizing Your Pandas Log Data Processing (Ditch the Slow Loops!)

Hey there! Let's swap that clunky loop for Pandas' native vectorized operations—they're way faster, cleaner, and built exactly for this kind of data task. Here's how to tackle your problem step by step:

First, let's start with your original data (I'll add a quick step to convert the time column to datetime—this is almost always a smart move for time-based log data):

import pandas as pd

# Your original log DataFrame
df1 = pd.DataFrame(
    [[1, 'confirmed', '01/01/2017 14:05:00'], [1, 'picked', '01/01/2017 14:10:00']],
    columns=['ID', 'log', 'time']
)

# Convert time column to datetime type (critical for any time-based operations later)
df1['time'] = pd.to_datetime(df1['time'])

Method 1: Attach Previous Row Data Directly, Then Filter

This is the simplest approach if you want to keep picked rows and their corresponding prior records in one structured DataFrame:

# Add columns to store the previous row's log entry and timestamp
df1['prev_log'] = df1['log'].shift(1)
df1['prev_time'] = df1['time'].shift(1)

# Filter to only keep rows where log is 'picked'—this gives you the picked entry AND its prior record
df2 = df1[df1['log'] == 'picked'].copy()

# Optional: Clean up df2 to only keep relevant columns
# df2 = df2[['ID', 'log', 'time', 'prev_log', 'prev_time']]

Method 2: Extract Picked Rows and Their Previous Rows Separately

If you need the prior rows as a distinct set (e.g., to compare side-by-side with picked entries), use index-based selection:

# Get indices of all rows where log is 'picked'
picked_indices = df1[df1['log'] == 'picked'].index

# Get indices of their previous rows (skip the first row to avoid index out-of-bounds errors)
prev_indices = [idx - 1 for idx in picked_indices if idx > 0]

# Extract the previous rows and picked rows, reset indices to align them
prev_rows = df1.loc[prev_indices].reset_index(drop=True)
picked_rows = df1.loc[picked_indices].reset_index(drop=True)

# Combine into df2 with clear labels for each record type
df2 = pd.concat(
    [prev_rows, picked_rows],
    axis=1,
    keys=['previous_record', 'picked_record']
)

Why This Beats Looping

  • Speed: Pandas' vectorized operations (like shift(), boolean indexing, and loc) run on optimized C code under the hood. For large datasets, this will be orders of magnitude faster than a Python loop.
  • Readability: The code clearly states your intent (filter picked rows, grab prior records) instead of getting lost in loop mechanics.
  • Bug Resistance: Chain indexing (like df1['log'][row]) often leads to unintended SettingWithCopy warnings or errors—using Pandas' official methods avoids this mess.

内容的提问来源于stack exchange,提问作者Hemingway_PL

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 04:22:41