You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在未排序DataFrame中按分组计算时间维度的金额百分比变化?

Solution for Calculating Percentage Change Without Sorting Original DataFrame

Got it, let's fix this issue for you. First, let's break down why your original code returned all NaNs: you grouped by ['Name', 'Date', 'ID'], which means every single row became its own group (since each combination of those three is unique in your data). With only one row per group, pct_change() can't calculate any differences—hence all NaNs.

The correct approach is to group by only Name and ID, sort each group internally by Date to get the chronological order, compute the percentage change, then map those results back to your original DataFrame while preserving its original row order.

Step-by-Step Code Implementation

First, let's make sure your Date column is in datetime format (critical for correct sorting):

import pandas as pd

# Your original data (for reference)
data = [
    ["Faye", "2019-12-31", "A", 2],
    ["Faye", "2020-03-01", "A", 3],
    ["Faye", "2019-09-30", "A", 1],
    ["Mike", "2019-09-30", "A", 7],
    ["Mike", "2019-12-31", "A", 7],
    ["Faye", "2019-09-30", "B", 10],
    ["Mike", "2019-12-31", "B", 12],
    ["Faye", "2019-12-31", "B", 8],
    ["Faye", "2019-06-30", "B", 5],
    ["Mike", "2019-09-30", "B", 10],
    ["Faye", "2019-09-30", "C", 5],
    ["Mike", "2018-03-31", "D", 5],
]
df = pd.DataFrame(data, columns=["Name", "Date", "ID", "Amount"])
df["Date"] = pd.to_datetime(df["Date"])  # Convert to datetime for proper sorting

Next, define a function to handle each group's calculation, then apply it while preserving the original order:

def compute_change(group):
    # Sort the group chronologically by Date (only internal to the group)
    sorted_group = group.sort_values("Date")
    # Calculate percentage change: (current - previous)/previous
    sorted_group["% Change"] = sorted_group["Amount"].pct_change()
    # Mark first entry in the group as "New"
    sorted_group["% Change"] = sorted_group["% Change"].apply(lambda x: "New" if pd.isna(x) else x)
    # Set to NaN if amount didn't change (pct_change = 0)
    sorted_group["% Change"] = sorted_group["% Change"].apply(lambda x: pd.NA if x == 0 else x)
    return sorted_group

# Apply the function to each (Name, ID) group, then restore original row order with sort_index()
result_df = df.groupby(["Name", "ID"], group_keys=False).apply(compute_change).sort_index()

Key Details

  • We only group by Name and ID to keep all related entries together.
  • Sorting happens only inside each group, so your original DataFrame's row order stays intact (we use sort_index() at the end to revert to the original order after grouping).
  • The logic handles all your requirements:
    • First entry in each group → marked as New
    • Amount unchanged → marked as NaN
    • Valid percentage changes → returned as the raw decimal value (multiply by 100 if you want percentage format, e.g., x * 100 in the lambda)

Example Output Snippet

For the first few rows of your original data, the result will look like this:

NameDateIDAmount% Change
0Faye2019-12-31A21.0
1Faye2020-03-01A30.5
2Faye2019-09-30A1New

Notice row 2 (the earliest entry in Faye's A group) is marked New, row 0 shows a 100% increase from row 2, and row 1 shows a 50% increase from row 0—all while keeping the original row order.

内容的提问来源于stack exchange,提问作者user53526356

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.07 15:18:13