You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python移除.dat文件指定列或替换空值的技术实现需求

解决方案:处理含未转义引号的CSV列问题

问题根源

原数据的line description字段包含未转义的双引号(如示例中的"LIFE LOGO"),不符合CSV规范(字段内引号需用双引号转义,即""LIFE LOGO""),导致csv.reader解析时错误拆分字段,引发数据移位。原代码仅全局替换引号,未修复解析错误,因此无法解决问题。

修正方案(两种可选)

方案1:移除line description列

以下代码先修复CSV格式错误,再精准移除目标列:

import os
import csv
import sys
from shutil import copyfile

sourcedir=sys.argv[1]
sourcefile=sys.argv[2]
targetdir=sys.argv[3]
targetfile=sys.argv[4]
qualifierIN=sys.argv[5]
delimiterIN=sys.argv[6]
qualifierOUT=sys.argv[7]
delimiterOUT=sys.argv[8]

curDir = os.getcwd()

# 检查源目录与文件
if not os.path.exists(sourcedir):
    print("Source Directory Not Found : ", sourcedir)
    sys.exit(5)
os.chdir(sourcedir)
if not os.path.isfile(sourcefile):
    print("Source File Not Found : ", sourcefile)
    sys.exit(5)
print("Source Directory :", sourcedir)
print("Source File :", sourcefile)

# 检查并创建目标目录
if not os.path.exists(targetdir):
    print("Target Directory Not Found :", targetdir)
    os.makedirs(targetdir)
print("Target Directory :", targetdir)

try:
    with open(sourcefile, 'rt', newline='') as fin:
        # 读取并处理表头,定位目标列索引
        header_line = next(fin).strip()
        header = [col.strip(qualifierIN) for col in header_line.split(delimiterIN)]
        try:
            line_desc_idx = header.index('line description')
        except ValueError:
            print("Column 'line description' not found in header")
            sys.exit(10)
        
        # 重置文件指针到开头
        fin.seek(0)
        processed_rows = []
        
        # 逐行修复未转义引号并解析
        for line in fin:
            line = line.rstrip('\n')
            in_quote = False
            processed_chars = []
            for char in line:
                if char == qualifierIN:
                    if in_quote:
                        # 引号内的引号转义为双引号
                        processed_chars.append(qualifierIN * 2)
                    else:
                        processed_chars.append(char)
                        in_quote = True
                else:
                    processed_chars.append(char)
            # 解析处理后的行
            processed_line = ''.join(processed_chars)
            row = next(csv.reader([processed_line], delimiter=delimiterIN, quotechar=qualifierIN))
            processed_rows.append(row)
        
        # 写入处理后的数据
        os.chdir(targetdir)
        with open(targetfile, 'wt', newline='') as fout:
            writer = csv.writer(fout, quoting=csv.QUOTE_ALL, quotechar=qualifierOUT, delimiter=delimiterOUT)
            
            for row in processed_rows:
                # 移除目标列
                if len(row) > line_desc_idx:
                    del row[line_desc_idx]
                
                # 处理字段内的特殊字符
                for idx in range(len(row)):
                    row[idx] = row[idx].replace(delimiterOUT, "")
                    row[idx] = row[idx].replace(qualifierOUT, "")
                    row[idx] = row[idx].replace(';', "")
                
                writer.writerow(row)
                
except csv.Error as e:
    print("CSV Error: ", str(e))
    sys.exit(25)
except Exception as e:
    print("Error : ", str(e))
    sys.exit(35)
finally:
    os.chdir(curDir)

# 替换源文件
try:
    os.chdir(sourcedir)
    os.remove(sourcefile)
    copyfile(os.path.join(targetdir, targetfile), sourcefile)
except Exception as e:
    print(str(e))
    sys.exit(45)

sys.exit(0)

方案2:置空line description列

只需将方案1中移除列的代码替换为置空逻辑:

# 置空目标列
if len(row) > line_desc_idx:
    row[line_desc_idx] = ""

关键修正点

  • 修复CSV格式错误:逐行处理未转义的内部引号,确保csv.reader能正确解析每个字段,避免拆分错误。
  • 精准定位目标列:通过表头找到line description的索引,确保操作的是正确列。
  • 避免数据移位:直接移除或置空目标列,而非全局修改字符,从根源解决移位问题。

内容的提问来源于stack exchange,提问作者Mythri

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.23 09:44:52