如何将Python导入的CSV格式GDP数据转换为时间序列?
Hey there! Let's fix that time series conversion issue you're having with your GDP data. The problem with your current code is that you're passing the entire DataFrame to the pd.Series() constructor instead of extracting the actual GDP growth values. Here's how to do it correctly:
Step 1: Understand the Data Structure
Your imported DataFrame has a single column GDP growth (%) with 12 rows of values. To create a time series, we need to pair these values with a proper datetime index.
Step 2: Correctly Create the Time Series Series
Instead of passing the whole DataFrame as data, extract the values from the GDP growth (%) column first. Also, make sure your datetime index length matches the number of data points (12 rows in your case):
import pandas as pd # Your existing import code GDP = pd.read_csv('GDP.csv') # Extract the GDP growth values from the DataFrame gdp_growth_values = GDP['GDP growth (%)'] # Create a datetime index that matches the number of data points # Using periods=len(gdp_growth_values) ensures we don't have a mismatch date_index = pd.date_range(start='01-2010', periods=len(gdp_growth_values), freq='Q') # Create the time series Series gdp_ts = pd.Series(data=gdp_growth_values.values, index=date_index)
Step 3: Alternative - Convert the Original DataFrame to a Time Series
If you prefer to keep the DataFrame structure (which is also valid for time series operations), you can simply set its index to the datetime range:
GDP.index = pd.date_range(start='01-2010', periods=len(GDP), freq='Q') # If you still want a Series later, just extract the column: gdp_ts = GDP['GDP growth (%)']
Why Your Original Code Failed
When you passed GDP (the entire DataFrame) to pd.Series(), pandas was trying to create a Series where each element was a row from the DataFrame, which is why you got those odd-looking entries like 2010-03-31 (G,.... By extracting the specific column values, you ensure the Series holds the GDP growth numbers paired with the correct datetime index.
Once you have this setup, you can perform all standard time series operations (like resampling, rolling averages, etc.) on gdp_ts or the indexed DataFrame.
内容的提问来源于stack exchange,提问作者Nick555

