从HUE HDFS读取文件时出现Python tokening错误如何解决?
问题根因
- 编码报错已经通过
cp1252配置解决,无需调整 - 后续的
ParserError是因为错误配置了sep=" "参数:pandas会按空格将每行内容拆分为多列,第一行拆分出7个字段,第二行拆分出16个字段,列数不匹配触发解析错误。你当前的数据集每行仅对应1条评价,无多列拆分需求,按空格拆分完全不符合数据读取逻辑。
可行解决方案
方案1:调整read_csv参数直接读取
无需拆分单列内容,指定按换行符分隔即可,代码如下:
with client.read(hdfsPath + "/test data/negative.csv") as reader: short_neg = pd.read_csv( reader, header = None, encoding = 'cp1252', sep = '\n', # 按换行符分隔,每行对应一条记录 engine = 'python' ) # 重命名列方便后续NLP处理 short_neg.columns = ['comment']
方案2:手动按行读取后转DataFrame
如果仍有解析异常,可以跳过pandas自动解析逻辑,直接按行读取后构造DataFrame,兼容性更强:
with client.read(hdfsPath + "/test data/negative.csv", encoding='cp1252') as reader: # 读取所有非空行,自动去除首尾空白符 comment_list = [line.strip() for line in reader if line.strip()] short_neg = pd.DataFrame(comment_list, columns=['comment'])
可选优化
如果仍存在零星编码报错,可以在读取时添加errors='ignore'配置,忽略无法解码的异常字符,不会影响整体情感分析效果:
encoding = 'cp1252', errors='ignore'
内容的提问来源于stack exchange,提问作者Tariq335
相关产品推荐
相关产品推荐

