Pandas新手求助:如何存储(x,y)对应n维列表至DataFrame并导出CSV?
Hey there! No worries at all—we all start somewhere with Pandas 😊 Let's break this down for you:
Can you store this data in a DataFrame and save it as CSV?
Absolutely, but you need to be aware of a key limitation of CSV files: they're flat, text-based formats. That means your n-dimensional lists (like P11) will get converted to string representations when saved to CSV. When you load the CSV back into Pandas, you'll have to parse those strings back into nested lists manually.
Here's a quick example to show you how this works:
import pandas as pd import ast # Sample nested data x_values = [["x1_val", "x2_val", "x3_val"]] y_data = { "y1": [[[1,2], [3,4]], [[5,6], [7,8]], [[9,10], [11,12]]], # P11, P12, P13 (each is 2D list) "y2": [[[13,14], [15,16]], [[17,18], [19,20]], [[21,22], [23,24]]], "y3": [[[25,26], [27,28]], [[29,30], [31,32]], [[33,34], [35,36]]] } # Create DataFrame df = pd.DataFrame({ "x1": [x_values[0][0]], "x2": [x_values[0][1]], "x3": [x_values[0][2]], "P11": [y_data["y1"][0]], "P12": [y_data["y1"][1]], "P13": [y_data["y1"][2]], "P21": [y_data["y2"][0]], "P22": [y_data["y2"][1]], "P23": [y_data["y2"][2]], "P31": [y_data["y3"][0]], "P32": [y_data["y3"][1]], "P33": [y_data["y3"][2]] }) # Save to CSV df.to_csv("nested_data.csv", index=False) # Load back from CSV and parse nested lists loaded_df = pd.read_csv("nested_data.csv") for col in loaded_df.columns: if col.startswith("P"): loaded_df[col] = loaded_df[col].apply(ast.literal_eval) # Now loaded_df['P11'] is back to a 2D list print(type(loaded_df['P11'][0])) # Output: <class 'list'>
Better storage alternatives for nested data
If you want to avoid parsing strings and preserve the nested structure directly, these options are more suitable:
1. Pickle/Joblib
These are Python-specific serialization formats that can save the exact structure of your DataFrame (including nested lists) without conversion. Joblib is especially good for large datasets.
# Save with Pickle df.to_pickle("nested_data.pkl") # Load back loaded_df_pkl = pd.read_pickle("nested_data.pkl") # Or use Joblib import joblib joblib.dump(df, "nested_data.joblib") loaded_df_joblib = joblib.load("nested_data.joblib")
2. HDF5
HDF5 is a binary format designed for large, complex datasets. Pandas has built-in support for it, and it preserves nested structures while being efficient for big data.
# Save to HDF5 df.to_hdf("nested_data.h5", key="data", mode="w") # Load back loaded_df_hdf = pd.read_hdf("nested_data.h5", key="data")
3. JSON
JSON is a human-readable format that handles nested structures better than CSV. Pandas can read/write JSON files, though it's less efficient than binary formats for very large data.
# Save to JSON df.to_json("nested_data.json", orient="records") # Load back loaded_df_json = pd.read_json("nested_data.json", orient="records")
Final recommendation
- Use CSV only if you need the data to be readable by non-Python tools and don't mind the extra parsing step.
- Use Pickle/Joblib for quick, easy storage when you'll only work with the data in Python.
- Use HDF5 if you're dealing with large datasets and need efficient storage/access.
- Use JSON for a balance between readability and nested structure support.
内容的提问来源于stack exchange,提问作者Abolfazl

