使用Pandas apply函数时DataFrame首行数据为何被重复打印两次?
Great question! The double print of your first row occurs because pandas runs your function twice on the first row: once as a "trial run" to infer the structure of your function's output, and again when applying the function to all rows in the DataFrame.
Here's a detailed breakdown:
- When you call
df.apply(func, axis=1), pandas needs to know what kind of output your function produces (a scalar, Series, list, etc.) to correctly construct the final result. - To figure this out, pandas first runs your function on the first row of the DataFrame. This is the first time you see the print output for row 0.
- Once pandas confirms the output type, it then runs the function on every row of the DataFrame—including the first row again. That’s why you see the print output for row 0 a second time.
- Rows 1 and 2 are only processed once during the full application, so their print outputs appear once each.
How to Prevent the Double Print
If you want to skip the trial run, explicitly tell pandas what output structure to expect using the result_type parameter. Since your function returns a Series (squaring each element of the input row), use result_type='expand':
df[["a","b"]].apply(func, axis=1, result_type='expand')
This tells pandas to expand the function’s output into a DataFrame, eliminating the need for the initial test run. Now your print statements will execute only once per row.
Content of the question originates from Stack Exchange, asked by Bhuvan M

