numpy列堆叠告警、CSV导出Excel异常及正则替换问题求解
NLP文本处理脚本三类问题修复方案
以下修复均基于原有代码逻辑实现,不改变业务输出结果。
问题1:np.column_stack触发VisibleDeprecationWarning
原因
numpy 1.24及以上版本禁止直接从长度不一致的不规则嵌套列表创建ndarray。给np.column_stack传dtype参数报错是因为该函数本身不接受这个入参,不需要指定dtype即可修复。
修复代码
替换原代码中merged_lists = np.column_stack([nouns, lemmas, poses, xposes, heads, deprels])为以下内容:
# 先将每列转为1D对象数组,规避numpy不规则序列检测 col_list = [nouns, lemmas, poses, xposes, heads, deprels] col_arrays = [np.array(col, dtype=object) for col in col_list] merged_lists = np.column_stack(col_arrays)
问题2:CSV双击用Excel打开排版错乱
原因
两个问题导致排版异常:
- 写CSV时使用无BOM的UTF-8编码,中文Windows环境下Excel默认用GBK编码打开UTF-8无BOM文件会出现识别错乱
- 原代码遍历行时固定取
len(string[0])作为列数,存在索引bug,且手动写入的多余空行会破坏CSV结构
修复代码
直接替换原to_csv函数为以下版本:
def to_csv(string, name): # 用utf-8-sig编码写入,自动添加BOM头让Excel直接识别编码 with open("CSV_" + str(name[:-4]) + ".csv", 'w', newline='', encoding='utf-8-sig') as c: writer = csv.writer(c, delimiter=',') for i in range(len(string)): # 修复原代码固定取第一行列数的索引bug for j in range(len(string[i])): writer.writerow(string[i][j])
问题3:re.sub特殊字符替换与原链式replace逻辑不一致
原因
之前写的正则存在三个问题:
- 字符组中
.-未做转义/位置处理,存在正则范围匹配歧义 - 遗漏了原逻辑中的
ℹ字符替换、eos转换行、\n删空行、连续4空格替换为单空格的步骤 - 替换顺序和原逻辑不匹配,导致空行未被清理,后续
splitlines读取到大量空行,连带影响emoji清理后的处理逻辑
修复代码
直接替换原create_filtered_text函数为以下版本,正则替换结果和原链式replace100%一致:
def create_filtered_text(string): x = open(string, encoding='utf-8') raw_text = x.read() x.close() text = raw_text.replace('eos', '\n') # 字符组内将-放在末尾,无需转义即可匹配字面量,覆盖所有原逻辑需要替换的单字符 text = re.sub(r'[\d,"·?¿:;!¡.ℹ-]', ' ', text) # 严格对齐原逻辑的替换顺序 text = text.replace(' \n', '') text = text.replace(' ', ' ') text = remove_emoji(text) x = open("FILTERED_" + string, 'w', encoding='utf-8') x.write(text) x.close()
内容的提问来源于stack exchange,提问作者Questioneer
相关产品推荐
相关产品推荐

