You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将类实例变量转为两列Pandas DataFrame?附tidypolars实现

Nice question! Let's walk through how to solve this with both Pandas (optimized for million-scale variables) and tidypolars.

Pandas Solution for Large-Scale Scenarios

When dealing with millions of instance variables, we need an efficient approach that avoids slow Python-level loops. The key here is leveraging Pandas' optimized underlying operations and the fact that a Python class instance stores its variables in the __dict__ attribute (a dictionary mapping variable names to their values).

Here's the step-by-step implementation:

import pandas as pd

class some_class:
    arg1 = 12345

    def __init__(self, arg1):
        self.variable1 = 1
        self.variable2 = 123
        self.variable_arg1 = arg1
        # Add your million-scale variables here...

# Create an instance of the class
obj = some_class(arg1=6789)

# Generate the two-column DataFrame efficiently
df = (
    pd.DataFrame.from_dict(obj.__dict__, orient='index', columns=['value'])
    .reset_index()
    .rename(columns={'index': 'variable'})
)

# Preview the result
print(df)

Why this works for large data:

  • obj.__dict__ directly gives us all instance variables as a key-value pair dictionary, no manual collection needed.
  • pd.DataFrame.from_dict uses optimized C-level operations to build the DataFrame, avoiding slow loops.
  • The chain of reset_index() and rename() is also vectorized, so it handles millions of rows without performance issues.
Tidypolars Solution

Tidypolars is a tidy interface for Polars, a high-performance dataframe library built in Rust. It's perfect for large-scale data thanks to its memory efficiency and speed. We can use Polars' built-in melt function to reshape our data into the desired two-column format.

Here's how to do it:

from tidypolars import DataFrame
import polars as pl

# Reuse the same class instance
obj = some_class(arg1=6789)

# Option 1: Use Polars' melt for optimal performance
df_tidy = DataFrame(
    pl.from_dict(obj.__dict__)
    .melt(variable_name='variable', value_name='value')
)

# Option 2: Directly construct from key-value pairs (simpler for small data, still efficient for large)
df_tidy = DataFrame([{'variable': k, 'value': v} for k, v in obj.__dict__.items()])

# Preview the result
print(df_tidy)

Notes for large-scale use:

  • Option 1 is preferred for millions of variables because Polars' melt is optimized for columnar data, making it faster and more memory-efficient than list comprehensions for huge datasets.
  • Tidypolars maintains the tidy syntax while leveraging Polars' Rust engine, so it outperforms Pandas in most large-data scenarios.

内容的提问来源于stack exchange,提问作者user10443249

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.04 10:22:54