You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

sum、'sum'与np.sum在Pandas df.agg方法中的底层差异及单引号含义探究

Differences Between sum, 'sum', and np.sum in Pandas agg()

Great question! Let’s break down exactly how each of these works under the hood, and clarify what the single-quoted string notation means in this context.

What does 'sum' (the string) do?

When you pass a string like 'sum' to agg(), pandas looks up its own optimized aggregation method for the data structure you’re working with (either a Series or DataFrame). Specifically, it maps to pandas.core.generic.NDFrame.sum, which is built explicitly for pandas objects. Key details:

  • It’s designed to handle pandas features like indexes, missing values (by default, it skips NaN with skipna=True), and axis-based calculations.
  • This is the most "pandas-native" option, and it’s what the official docs prefer because it plays nicely with pandas’ data structures and expected behavior.

How does the Python built-in sum differ?

The bare sum here refers to Python’s native sum() function. While it will technically work on a pandas Series (since Series are iterable), it’s not optimized for pandas data:

  • It doesn’t handle NaN gracefully: if your Series contains any NaN, Python’s sum will return NaN (unlike pandas’ sum, which skips them by default).
  • It iterates over the Series values one by one, which is slower than pandas’ vectorized operations that leverage numpy under the hood.
  • It doesn’t understand pandas-specific parameters like axis, so it’s less flexible for DataFrame-wide aggregations.

What about np.sum?

np.sum is numpy’s vectorized sum function, which operates on numpy arrays (the underlying storage for most pandas Series). Here’s how it compares:

  • Like Python’s sum, np.sum will return NaN if the input array contains any NaN (unless you explicitly set skipna=True—note that numpy’s skipna support is limited to newer versions and specific functions).
  • It’s faster than Python’s native sum because it uses vectorized operations, but it lacks pandas-specific logic: it doesn’t preserve index context or handle pandas’ edge cases as seamlessly as the pandas-native 'sum'.
  • Pandas will pass the underlying numpy array of the Series to np.sum, so it computes the sum of the raw values, but doesn’t account for pandas’ default behavior like skipping NaN unless you specify it.

Quick Example of Behavioral Differences

Suppose you have a Series with a NaN:

import pandas as pd
import numpy as np
s = pd.Series([1, 2, np.nan])

print(s.agg('sum'))  # Output: 3.0 (skips NaN by default)
print(s.agg(sum))    # Output: nan (Python sum doesn't skip NaN)
print(s.agg(np.sum)) # Output: nan (numpy sum doesn't skip NaN by default)

This shows that while the outputs might look identical for clean data, they diverge significantly when missing values are present.

内容的提问来源于stack exchange,提问作者nutty about natty

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.27 21:34:05