如何在Python中显示Polars DataFrame的完整列宽?
如何完整显示Polars DataFrame的长文本列
问题背景
有如下包含长文本列的Polars DataFrame:
import polars as pl df = pl.DataFrame({ 'column_1': ['TF-IDF embeddings are done on the initial corpus, with no additional N-Gram representations or further preprocessing', 'In the eager API, the expression is evaluated immediately. The eager API produces results immediately after execution, similar to pandas. The lazy API is similar to Spark, where a plan is formed upon execution of a query, but the plan does not actually access the data until the collect method is called to execute the query in parallel across all CPU cores. In simple terms: Lazy execution means that an expression is not immediately evaluated.'], 'column_2': ['Document clusterings may misrepresent the visualization of document clusterings due to dimensionality reduction (visualization is pleasing for its own sake - rather than for prediction/inference)', 'Polars has two APIs, eager and lazy. In the eager API, the expression is evaluated immediately. The eager API produces results immediately after execution, similar to pandas. The lazy API is similar to Spark, where a plan is formed upon execution of a query, but the plan does not actually access the data until the collect method is called to execute the query in parallel across all CPU cores. In simple terms: Lazy execution means that an expression is not immediately evaluated.'] })
尝试以下配置后,文本仍被截断:
pl.Config.set_fmt_str_lengths = 200 pl.Config.set_tbl_width_chars = 200
显示结果:
shape: (2, 2) ┌───────────────────────────────────┬───────────────────────────────────┐ │ column_1 ┆ column_2 │ │ --- ┆ --- │ │ str ┆ str │ ╞═══════════════════════════════════╪═══════════════════════════════════╡ │ TF-IDF embeddings are done on th… ┆ Document clusterings may misrepr… │ │ In the eager API, the expression… ┆ Polars has two APIs, eager and l… │ └───────────────────────────────────┴───────────────────────────────────┘
解决方案
1. 修正配置调用方式
你之前的写法有误,Polars的Config参数需要通过方法调用设置,而非直接赋值。
2. 全局设置完整显示
通过设置足够大的字符长度,或传入None禁用文本截断:
# 禁用字符串长度截断,显示完整文本 pl.Config.set_fmt_str_lengths(None) # 设置表格宽度为足够大的值,避免列被压缩 pl.Config.set_tbl_width_chars(1000) # 可选:显示所有行,避免行截断 pl.Config.set_fmt_max_rows(None)
执行上述配置后再打印df,即可看到完整的长文本列。
3. 临时生效的上下文管理器
若不想修改全局配置,可使用上下文管理器临时应用设置:
with pl.Config(fmt_str_lengths=None, tbl_width_chars=1000, fmt_max_rows=None): print(df)
效果验证
配置生效后,DataFrame会完整显示所有文本内容,不再出现末尾的…截断标记。
内容的提问来源于stack exchange,提问作者Ahmad
相关产品推荐
相关产品推荐

