You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何确保多模块中Polars DataFrame共用同一字符串缓存?

解决Polars多模块Categorical列跨缓存Join的问题

问题背景

在多模块结构的Polars项目中,若模块在导入阶段就创建了带pl.Category类型的DataFrame,后续开启全局字符串缓存后,跨模块Join这些DataFrame时会触发错误:

exceptions.ComputeError: joins/or comparisons on categoricals can only happen if they were created under the same global string cache

这是因为模块导入时(早于pl.toggle_string_cache(True)执行)创建的Categorical列,与缓存开启后操作的DataFrame不在同一缓存上下文。直接调整导入顺序违反“import放在代码开头”的规范,需要更优雅的解决方案。

优雅解决方案

方案1:延迟加载DataFrame(推荐)

将模块内的DataFrame创建逻辑封装为函数,避免模块导入时自动初始化,而是在缓存开启后主动调用生成。

修改A.py:

import polars as pl

def get_df_A():
    return pl.DataFrame({
        'A': pl.Series(['a', 'b', 'c'], dtype=pl.Category)
    })

修改B.py:

import polars as pl

def get_df_B():
    return pl.DataFrame({
        'A': pl.Series(['a', 'b', 'c'], dtype=pl.Category),
        'B': pl.Series(['d', 'e', 'f'], dtype=pl.Category)
    })

__init__.py保持import在开头,开启缓存后生成DataFrame:

import polars as pl
from . import A, B

pl.toggle_string_cache(True)

df_A = A.get_df_A()
df_B = B.get_df_B()
df_C = df_A.join(df_B, on='A')

方案2:模块级初始化函数

如果需要模块内保留全局DataFrame变量,可通过初始化函数控制创建时机。

修改A.py:

import polars as pl

df_A = None

def init():
    global df_A
    df_A = pl.DataFrame({
        'A': pl.Series(['a', 'b', 'c'], dtype=pl.Category)
    })

修改B.py:

import polars as pl

df_B = None

def init():
    global df_B
    df_B = pl.DataFrame({
        'A': pl.Series(['a', 'b', 'c'], dtype=pl.Category),
        'B': pl.Series(['d', 'e', 'f'], dtype=pl.Category)
    })

__init__.py中开启缓存后触发模块初始化:

import polars as pl
from . import A, B

pl.toggle_string_cache(True)
A.init()
B.init()

df_C = A.df_A.join(B.df_B, on='A')

核心逻辑

所有需要参与跨模块Join/比较的pl.Category列,必须在同一全局字符串缓存开启的上下文中创建。通过延迟初始化或主动触发初始化,既能遵守import规范,又能保证缓存上下文的一致性。

内容的提问来源于stack exchange,提问作者NedDasty

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.25 11:58:08