如何用Pandas读取含跨行Numpy数组的CSV文件?
解决CSV跨行数组字段的读取问题
问题根源
你的CSV里lyrics_bert_embeddings字段是用np.array2string()生成的数组字符串,虽然用双引号包裹,但内容跨了多行。标准CSV解析器默认把换行当成新行的起始,自然会触发解析错误。
自动处理方案
不用手动修改CSV文件,这两种方法可以自动合并跨行内容:
1. 用Python的csv模块自定义合并逻辑
利用csv.reader的引号识别功能,把属于同一字段的多行内容拼接回去:
import csv import pandas as pd rows = [] with open('你的文件名.csv', 'r', newline='', encoding='utf-8') as f: reader = csv.reader(f, quoting=csv.QUOTE_ALL) header = next(reader) current_row = None for parts in reader: if not current_row: current_row = parts else: # 检查上一行最后一个字段是否是未闭合的引号内容 last_field = current_row[-1] if last_field.startswith('"') and not last_field.endswith('"'): # 把当前行内容拼接到上一行的最后一个字段 current_row[-1] += '\n' + ' '.join(parts) # 拼接后如果引号闭合了,就把这行存入列表 if current_row[-1].endswith('"'): rows.append(current_row) current_row = None else: rows.append(current_row) current_row = parts # 处理最后一行未完成的内容 if current_row: rows.append(current_row) # 转成DataFrame df = pd.DataFrame(rows[1:], columns=rows[0]) # 可选:把字符串格式的数组转成numpy数组 import numpy as np df['lyrics_bert_embeddings'] = df['lyrics_bert_embeddings'].apply(lambda x: np.fromstring(x.strip('"[]'), sep=' '))
2. 直接用pandas的read_csv参数配置
给Python引擎指定引号规则,让它正确识别跨行的引号包裹字段:
import pandas as pd import numpy as np df = pd.read_csv('你的文件名.csv', engine='python', quotechar='"', quoting=pd.io.common.csv.QUOTE_ALL, skipinitialspace=True) # 把字符串格式的数组转成numpy数组 df['lyrics_bert_embeddings'] = df['lyrics_bert_embeddings'].apply(lambda x: np.fromstring(x.strip('"[]'), sep=' '))
核心是quoting=pd.io.common.csv.QUOTE_ALL这个参数,它告诉解析器所有字段都用引号包裹,引号内的跨行内容属于同一个字段。
后续避坑建议
以后生成CSV时,别直接用np.array2string()输出数组,改成np.array2string(arr, separator=',')把数组转成单行的逗号分隔字符串,或者直接存储扁平化的列表,这样CSV解析就不会出现跨行问题了。
内容的提问来源于stack exchange,提问作者Fenrir
相关产品推荐
相关产品推荐

