You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何对两个DataFrame做减法获取数值差异及每日聚合DataFrame相关问题

Hey there! Let's tackle your two pandas DataFrame needs with practical, actionable examples—since you've got both a general subtraction request and a specific daily aggregated data use case.


1. General DataFrame Subtraction for Value Differences

First, for your core request to subtract two DataFrames and get corresponding value differences, here's how to approach it effectively:

Key Prerequisites

  • Your two DataFrames should have matching indices and column names (or at least align on the specific rows/columns you want to compare).
  • Pandas automatically aligns data by row/column labels, so mismatched labels will return NaN for those positions unless you specify a fill value.

Simple Direct Subtraction

The most straightforward method uses the - operator directly—this works perfectly if your DataFrames have identical structures:

# Example: df1 and df2 share the same index and columns
difference_df = df1 - df2

Flexible Subtraction with sub()

For more control (like handling missing values or partial column matches), use the DataFrame.sub() method. This lets you fill gaps in either DataFrame before subtraction:

# Fill missing values with 0 to avoid NaN results
difference_df = df1.sub(df2, fill_value=0)

# Subtract only specific target columns
target_columns = ['cpu_mean', 'mem_min', 'latency_std']
difference_df = df1[target_columns].sub(df2[target_columns], fill_value=0)

2. Subtraction for Daily Aggregated DataFrames

Now, for your specific case with daily aggregated DataFrames (like the 2018-02-11 dataset with metrics such as cpu, mem, load, each tracked via mean, min, max, std), here's how to compute meaningful differences between two such daily datasets:

Step 1: Ensure Consistent Structure

First, confirm both daily DataFrames have identical column names (e.g., cpu_mean, mem_max, gw_latency_std). This guarantees pandas aligns the correct stats for each metric.

Step 2: Compute Metric-by-Metric Differences

Use the same subtraction logic, but tailor it to your aggregated structure:

# Example: df_20180211 is your 2018-02-11 aggregated DataFrame
# df_20180210 is the aggregated DataFrame for the previous day
daily_diff = df_20180211 - df_20180210

# Filter to focus on specific metrics (e.g., cpu, mem, latency)
focus_metrics = ['cpu', 'mem', 'latency']
# Grab all columns related to these metrics
focus_columns = [col for col in df_20180211.columns if any(metric in col for metric in focus_metrics)]
focus_diff = df_20180211[focus_columns] - df_20180210[focus_columns]

Step 3: Interpret Results

  • Positive values mean the 2018-02-11 stat is higher than the comparison day.
  • Negative values mean the stat is lower than the comparison day.
  • Use fill_value=0 in sub() if you want to treat missing metrics as 0 instead of returning NaN.

内容的提问来源于stack exchange,提问作者Souvik Ray

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 04:16:43