Pandas新增列类型问题:event_id为何转为float及解决方法
Great question! Let's break down why your event_id values (180, 181) are showing up as floats (180.0, 181.0) and how to fix it.
Root Cause
The issue comes down to how pandas handles dynamic column addition and type inference with empty DataFrames:
- You start with
data_for_csv = pd.DataFrame(columns=["ds", "yhat", "yhat_lower", "yhat_upper"]), an empty DataFrame that doesn't include theevent_idcolumn. - When you append rows containing the
event_idcolumn, pandas has to add this column on the fly. - Since the initial DataFrame was empty, pandas prioritizes flexibility over strict typing: it defaults to float for the new column because float can accommodate missing values (NaN) whereas integer types cannot. Even though your
event_idvalues are integers, pandas assumes the column might eventually have missing data and converts it to float as a precaution.
In some cases, if other columns in your forecast DataFrame are float types, this can reinforce pandas' decision to cast event_id to float to maintain consistency, but this is secondary to the empty DataFrame issue.
Fixes to Keep event_id as Integer
1. Define event_id upfront with explicit dtype
Create your initial DataFrame with the event_id column included and specify its integer type from the start. This tells pandas exactly what type to use, eliminating implicit casting:
# Include event_id in columns and set its dtype explicitly data_for_csv = pd.DataFrame( columns=["event_id", "ds", "yhat", "yhat_lower", "yhat_upper"], dtype={"event_id": int} )
2. Convert the column to integer after appending
If you want to keep your original initial DataFrame setup, you can cast the event_id column back to integer once all appends are complete. Just ensure there are no missing values (NaN) in the column first (which there aren't in your case):
# After all appends finish data_for_csv["event_id"] = data_for_csv["event_id"].astype(int)
3. Enforce integer type when adding event_id to forecast
When adding the event_id column to the forecast DataFrame, explicitly set its type to integer to avoid any pre-append casting:
# Explicitly create an integer series for event_id forecast["event_id"] = pd.Series([event_id] * len(forecast), dtype=int)
Verification
After implementing any fix, check the column types to confirm event_id is now integer:
print(data_for_csv.dtypes)
You should see output like this:
event_id int64 ds object yhat float64 yhat_lower float64 yhat_upper float64 dtype: object
内容的提问来源于stack exchange,提问作者Joey Coder

