如何将含104个pandas DataFrame的Python字典转为R可读取的DataFrame列表
Got it, saving a Python dictionary with 104 pandas DataFrames so R can read it as a list of DataFrames is totally solvable—JSON often falls short here because it struggles with preserving tabular data types and handling nested structures cleanly. Let’s dive into robust, easy-to-use alternatives:
Option 1: Feather Format (Lightweight & Fast)
Feather was built specifically for cross-language data exchange, so it preserves data types perfectly and is blazingly fast. It’s my go-to for small-to-medium datasets like yours.
Python: Save the DataFrames
First, install the required package:
pip install feather-format
Then save each DataFrame to a dedicated Feather file:
import pandas as pd from pathlib import Path # Replace with your actual dictionary of DataFrames df_dict = {"df_01": pd.DataFrame(...), "df_02": pd.DataFrame(...), ...} # Create a folder to store all Feather files output_folder = Path("r_compatible_data") output_folder.mkdir(exist_ok=True) # Loop through the dictionary and save each DataFrame for df_name, df in df_dict.items(): df.to_feather(output_folder / f"{df_name}.feather")
R: Read into a DataFrame List
Install the arrow package (it handles Feather files seamlessly):
install.packages("arrow")
Then load all files into a named list:
library(arrow) # Get paths to all Feather files feather_files <- list.files("r_compatible_data", pattern = "\\.feather$", full.names = TRUE) # Extract file names to use as list names names(feather_files) <- tools::file_path_sans_ext(basename(feather_files)) # Read all files into a list of DataFrames df_list <- lapply(feather_files, read_feather) # Check the result str(df_list)
Option 2: R Native RData Format (Direct Conversion)
If you want to skip file-per-DataFrame and directly save an R-compatible list, use rpy2 to convert the Python dictionary to an R list and save it as an RData file. This is the most direct approach.
Python: Convert & Save
Install rpy2 first:
pip install rpy2
Then convert and save:
import pandas as pd import rpy2.robjects as ro from rpy2.robjects import pandas2ri from rpy2.robjects.packages import importr # Enable conversion between pandas DataFrames and R data.frames pandas2ri.activate() # Your dictionary of DataFrames df_dict = {"df_01": pd.DataFrame(...), "df_02": pd.DataFrame(...), ...} # Convert Python dict to R ListVector (each element is an R data.frame) r_data_list = ro.ListVector({name: pandas2ri.py2rpy(df) for name, df in df_dict.items()}) # Save the R list to an RData file base = importr("base") base.save(r_data_list, file="data_frame_list.RData")
R: Load the List Directly
Just load the file, and you’ll have the list ready to use:
# Load the RData file load("data_frame_list.RData") # The list is named r_data_list (from the Python code) str(r_data_list)
Option 3: Parquet Format (Great for Large/Compressed Data)
If your DataFrames are large and you want efficient compression, Parquet is ideal. It’s a columnar storage format that works seamlessly across Python and R.
Python: Save as Parquet
Install pyarrow (handles Parquet):
pip install pyarrow
Save each DataFrame:
import pandas as pd import pyarrow as pa import pyarrow.parquet as pq from pathlib import Path df_dict = {"df_01": pd.DataFrame(...), "df_02": pd.DataFrame(...), ...} output_folder = Path("parquet_data") output_folder.mkdir(exist_ok=True) for df_name, df in df_dict.items(): # Convert pandas DataFrame to Arrow Table table = pa.Table.from_pandas(df) # Write to Parquet file pq.write_table(table, output_folder / f"{df_name}.parquet")
R: Read into a List
Use the arrow package again:
library(arrow) parquet_files <- list.files("parquet_data", pattern = "\\.parquet$", full.names = TRUE) names(parquet_files) <- tools::file_path_sans_ext(basename(parquet_files)) df_list <- lapply(parquet_files, read_parquet)
Why JSON Was a Problem
JSON is a text-based format that doesn’t natively support tabular data types (like dates, integers, or categorical variables)—it often converts these to strings, forcing you to re-cast them in R. Additionally, parsing a nested JSON structure (a dict of DataFrames) into an R list of data.frames requires extra, messy processing. The formats above avoid all that hassle.
内容的提问来源于stack exchange,提问作者MadmanLee

