读取UTF-16 LE编码CSV后,Pandas数据数值计算问题求助
解决UTF-16 LE编码CSV读取后数值列无法转为浮点型的问题
问题根源
你读取的UTF-16 LE编码CSV中,数值列表面显示正常,但实际包含**不可见的空字节(\x00)**或多余空白字符,导致Pandas将其识别为object类型(字符串)而非数值类型,触发can't multiply sequence by non-int of type 'float'错误。
解决方案
方案1:读取时直接转换数值列
在pd.read_csv中通过converters参数指定转换规则,提前清理字符并转为浮点型:
import pandas as pd columns = ["date","time","ch104","alarm104","ch114","alarm114","ch115","alarm115","ch116","alarm116","ch117","alarm117","ch118","alarm118"] # 定义清理并转换的函数 def clean_to_float(s): cleaned = str(s).strip().replace('\x00', '') # 移除空字节和首尾空白 return float(cleaned) if cleaned else pd.NA df = pd.read_csv( cal_file, sep='[, ]', encoding='UTF-16 LE', names=columns, header=15, on_bad_lines='skip', engine='python', converters={ 'ch104': clean_to_float, 'ch114': clean_to_float, 'ch115': clean_to_float, 'ch116': clean_to_float, 'ch117': clean_to_float, 'ch118': clean_to_float } ) ch104 = df['ch104'].dropna() # 验证计算 print(ch104 * 1.2)
方案2:读取后批量处理列
如果已经完成文件读取,可对所有ch开头的列批量清理转换:
# 遍历所有数值列 for col in df.columns: if col.startswith('ch'): # 清理空字节、空白,转为浮点型 df[col] = ( df[col] .astype(str) .str.strip() .str.replace('\x00', '') .replace('', pd.NA) .astype(float) ) ch104 = df['ch104'].dropna() # 执行数值计算 print(ch104.mean())
关键说明
UTF-16 LE编码的文件每个字符占用2字节,部分数值字符串会残留空字节\x00,这些字符肉眼不可见但会干扰Pandas的类型识别。通过str.replace('\x00', '')彻底移除空字节,再转为浮点型即可解决问题。
内容的提问来源于stack exchange,提问作者arv24_
相关产品推荐
相关产品推荐

