Polars读取CSV报错:分隔符§应为单字节字符却为2字节
解决Polars读取§作为分隔符的CSV文件报错问题
问题原因
§在UTF-8编码下是2字节字符(二进制为0xC2 0xA7),而Polars默认的Rust解析引擎仅支持单字节分隔符,因此触发"seperator ="§" should be a single byte character, but is 2 bytes long"错误。
解决方案
方案1:切换到Python解析引擎
Polars提供的Python引擎支持多字节分隔符,只需在read_csv中添加engine="python"参数即可:
import polars as po po_df = po.read_csv( "<full file path>", separator='§', has_header=True, quote_char='"', encoding='utf8', engine="python" ) print(po_df.columns)
注意:Python引擎速度略慢于默认Rust引擎,但能直接处理多字节分隔符,适合大多数场景。
方案2:替换分隔符为单字节字符
先手动读取文件内容,将§替换为一个不会出现在数据中的单字节字符(比如|),再用默认引擎解析:
import polars as po # 读取文件并替换分隔符 with open("<full file path>", 'r', encoding='utf8') as f: csv_content = f.read().replace('§', '|') # 从字符串加载数据到Polars po_df = po.read_csv( csv_content, separator='|', has_header=True, quote_char='"', encoding='utf8' ) print(po_df.columns)
注意:需确保替换的单字节字符未在原始数据中出现,否则会导致列解析错误。该方法保留了Rust引擎的高性能,适合大文件场景。
内容的提问来源于stack exchange,提问作者Kaustav Nandy
相关产品推荐
相关产品推荐

