如何将Python循环输出的分布拟合结果转换为DataFrame格式?
Solution to Convert Print Output to DataFrame
Here's how you can tweak your existing code to gather all results into a structured DataFrame instead of printing them line by line:
Modified Code
import pandas as pd def distribution_selection(csv_file=None, product_column=None, demand=None): # Read the CSV into a DataFrame (fixes the undefined df_all reference in your original code) df_all = pd.read_csv(csv_file) if csv_file else pd.DataFrame() results = [] # Store each product's data as a dictionary here for num in df_all[product_column]: # ... (keep your existing logic to compute best_fit, likelihoods, and parameters like mean, std, p_nbinom, r_binom, lambda_) # Start with common fields for every product product_result = { "Products": num, "BestFit": best_fit, "Likelihood": likelihoods[best_fit] } # Add parameters based on the best-fit distribution if best_fit == "norm": product_result["ParameterA"] = mean product_result["ParameterB"] = std elif best_fit == "nbinom": product_result["ParameterA"] = p_nbinom product_result["ParameterB"] = r_binom elif best_fit == "poisson": product_result["ParameterA"] = lambda_ product_result["ParameterB"] = pd.NA # Mark missing second parameter results.append(product_result) # Convert the list of dictionaries to a DataFrame output_df = pd.DataFrame(results) # Reorder columns to match your desired structure output_df = output_df[["Products", "BestFit", "Likelihood", "ParameterA", "ParameterB"]] # Print the final DataFrame (or return it for further use) print(output_df) return output_df
Key Changes Breakdown
- Result Collection: Instead of printing immediately, we store each product's data in a dictionary and add it to a list. This makes converting to a DataFrame straightforward.
- Parameter Handling: For each distribution type, we map the computed parameters to
ParameterAandParameterB. For Poisson (which only has one parameter), we usepd.NAto represent the missing value, which shows up asNAin the final output. - CSV Reading: Added a line to load the CSV file into
df_allsince your original code referenced this variable but didn't define it in the function. - Column Ordering: Reordered the DataFrame columns to exactly match the structure you requested.
Example Output
When you run the function, you'll get a neatly formatted DataFrame like this:
| Products | BestFit | Likelihood | ParameterA | ParameterB |
|---|---|---|---|---|
| 001.001 | nbinom | 6.317496e-15 | 0.002660521439486909 | 0.41659311972644725 |
| 001.002 | nbinom | 5.902081e-18 | 0.005335820123825622 | 0.7249662663271113 |
| 001.003 | nbinom | 2.871871e-13 | 0.00743701201046538 | 0.45081292375812926 |
| 001.004 | poisson | 0.0002870492567273848 | 15.333333333333334 |
内容的提问来源于stack exchange,提问作者Salih
相关产品推荐
相关产品推荐

