如何将类实例变量转为两列Pandas DataFrame?附tidypolars实现
Nice question! Let's walk through how to solve this with both Pandas (optimized for million-scale variables) and tidypolars.
When dealing with millions of instance variables, we need an efficient approach that avoids slow Python-level loops. The key here is leveraging Pandas' optimized underlying operations and the fact that a Python class instance stores its variables in the __dict__ attribute (a dictionary mapping variable names to their values).
Here's the step-by-step implementation:
import pandas as pd class some_class: arg1 = 12345 def __init__(self, arg1): self.variable1 = 1 self.variable2 = 123 self.variable_arg1 = arg1 # Add your million-scale variables here... # Create an instance of the class obj = some_class(arg1=6789) # Generate the two-column DataFrame efficiently df = ( pd.DataFrame.from_dict(obj.__dict__, orient='index', columns=['value']) .reset_index() .rename(columns={'index': 'variable'}) ) # Preview the result print(df)
Why this works for large data:
obj.__dict__directly gives us all instance variables as a key-value pair dictionary, no manual collection needed.pd.DataFrame.from_dictuses optimized C-level operations to build the DataFrame, avoiding slow loops.- The chain of
reset_index()andrename()is also vectorized, so it handles millions of rows without performance issues.
Tidypolars is a tidy interface for Polars, a high-performance dataframe library built in Rust. It's perfect for large-scale data thanks to its memory efficiency and speed. We can use Polars' built-in melt function to reshape our data into the desired two-column format.
Here's how to do it:
from tidypolars import DataFrame import polars as pl # Reuse the same class instance obj = some_class(arg1=6789) # Option 1: Use Polars' melt for optimal performance df_tidy = DataFrame( pl.from_dict(obj.__dict__) .melt(variable_name='variable', value_name='value') ) # Option 2: Directly construct from key-value pairs (simpler for small data, still efficient for large) df_tidy = DataFrame([{'variable': k, 'value': v} for k, v in obj.__dict__.items()]) # Preview the result print(df_tidy)
Notes for large-scale use:
- Option 1 is preferred for millions of variables because Polars'
meltis optimized for columnar data, making it faster and more memory-efficient than list comprehensions for huge datasets. - Tidypolars maintains the tidy syntax while leveraging Polars' Rust engine, so it outperforms Pandas in most large-data scenarios.
内容的提问来源于stack exchange,提问作者user10443249

