You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python将文本CSV拆分为含customer_id、date、comments的结构化CSV

解决CSV格式整理问题的Python方案

核心思路

  • 读取原CSV时不按逗号分割,先将整行内容作为字符串处理
  • 从每行开头提取前两个有效字段:客户ID(数字)、日期(YYYY-MM-DD格式)
  • 将剩余所有内容拼接作为评论字段
  • 最后把整理好的数据写入结构化CSV文件

代码实现

import csv

# 替换为你的文件路径
input_path = "your_input.csv"
output_path = "structured_output.csv"

with open(input_path, 'r', encoding='utf-8') as infile, open(output_path, 'w', newline='', encoding='utf-8') as outfile:
    # 初始化CSV写入器,指定目标字段
    writer = csv.DictWriter(outfile, fieldnames=['customer_id', 'date', 'comments'])
    writer.writeheader()
    
    for line in infile:
        line = line.strip()
        if not line:
            continue
        
        # 拆分出客户ID、日期,剩余部分作为评论主体
        split_line = line.split(maxsplit=2)
        customer_id = split_line[0].strip('"')  # 去除开头的引号
        date = split_line[1]
        # 拼接所有评论内容,处理原数据里的引号和逗号
        comments_part = split_line[2].strip('"') + line[len(split_line[0]+split_line[1]+split_line[2]):]
        comments = comments_part.replace('", ', ', ').rstrip(',')
        
        # 写入整理后的行
        writer.writerow({
            'customer_id': customer_id,
            'date': date,
            'comments': comments
        })

代码说明

  • split(maxsplit=2):仅分割前两个空格,确保客户ID和日期被精准提取,剩余内容全部归为评论
  • 引号处理:原数据开头字段带双引号,用strip('"')清除
  • 评论拼接:合并分割后剩余部分与原行后续内容,保留原评论的标点和格式
  • 用DictWriter写入,保证字段与目标结构完全对应

示例验证

针对你给出的示例行:

"216604 2022-08-22 Overall", this bank is satisfactory.,,,

处理后生成的结构化字段:

  • customer_id: 216604
  • date: 2022-08-22
  • comments: Overall, this bank is satisfactory.

内容的提问来源于stack exchange,提问作者Paco

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.22 15:50:11