如何在Python中从ARIMA模型摘要提取指定值并存入DataFrame
I want to extract specific values (such as Model and AIC) from the ARIMA model results summary and save them into a DataFrame. After running
result.summary(), I get the following ARIMA model results table:ARIMA Model Results
Dep. Variable: D.Order_Count
No. Observations: 543
Model: ARIMA(4, 1, 4)
Log Likelihood -1589.811
Method: css-mle
S.D. of innovations 4.492
Date: Wed, 11 Apr 2018
AIC 3199.622
Time: 12:35:05
BIC 3242.593
Sample: 07-02-2015 - 07-31-2017
HQIC 3216.424How can I achieve this in Python?
Answer:
Got it, this is a super common task, and the most reliable way to do it is by using the built-in properties of the statsmodels ARIMA result object—no need to mess around parsing the text summary (which can break if the summary format changes between library versions). Here's a straightforward step-by-step solution:
First, make sure you have pandas and statsmodels installed. Then, assuming you've already fitted your ARIMA model and have the result object ready, follow these steps:
- Collect metrics using result object properties
Statsmodels stores every value shown in the summary as a direct attribute of the result object. We can gather them into a dictionary first:
import pandas as pd # Assume 'result' is your fitted ARIMA model result object extracted_stats = { "Model": f"ARIMA{result.model.order}", # Grabs the (p,d,q) tuple and formats it to match the summary "No. Observations": result.nobs, "Log Likelihood": result.llf, "Method": result.method, "S.D. of innovations": result.sigma2 ** 0.5, # sigma2 is the variance; square root gives standard deviation "AIC": result.aic, "BIC": result.bic, "HQIC": result.hqic, "Sample": f"{result.data.dates.min().strftime('%m-%d-%Y')} - {result.data.dates.max().strftime('%m-%d-%Y')}" }
- Convert the dictionary to a DataFrame
Turn the dictionary into a single-row DataFrame (you can easily expand this later to compare multiple models by appending more dictionaries):
stats_df = pd.DataFrame([extracted_stats]) # Print the result or save to a file print(stats_df) # To save as CSV: stats_df.to_csv("arima_model_metrics.csv", index=False)
Why this is better than parsing text:
Directly accessing the result object's properties is way more stable than using regex to scrape the summary text. The summary's formatting might shift slightly between statsmodels updates, but the underlying properties will stay consistent.
If you really need to parse the summary text (not recommended):
For edge cases where you have to work with the raw text output, you can use regular expressions. Here's a quick example:
import re # Get the summary as plain text summary_text = result.summary().as_text() # Extract Model name model_match = re.search(r"Model:\s*(.*)", summary_text) model = model_match.group(1) if model_match else None # Extract AIC value aic_match = re.search(r"AIC\s*(.*)", summary_text) aic = float(aic_match.group(1)) if aic_match else None # Build dictionary and convert to DataFrame text_extracted = {"Model": model, "AIC": aic} text_df = pd.DataFrame([text_extracted])
But again, stick to the property method whenever possible—it's cleaner, faster, and less prone to errors.
内容的提问来源于stack exchange,提问作者anant saxena

