You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python读取损坏CSV:如何合并文本并添加双引号修复

问题描述

CSV文件存在格式损坏,commentText字段未被双引号包裹,导致该字段的多行文本被错误分割成多行记录。使用标准csv.DictReader读取时,只能获取到字段的开头部分([p])。

损坏的CSV示例

ProcessId,nid,CreatedDate,name,uid,Forum,cid,pid,hostname,Change,status,Thread,level,CommentOrder,bundle,deleted,lang,delta,Length_comment_text,field_data_comment_body,topic,isAnswer,CreatedAt,Updatedat,commentText,sticky,CommentRating,CommentRateTimes,Helpfull
1,1031762,2018-01-01 03:42:53,vimo,96977,LoRa®,1031762,0,204.2.166.150,2018-01-01 03:42:53,0,1031762,0,1031762,,0,und,0,1514,comment,How can RN2483 reach -146dBm sensitivity...,0,2018-01-01 03:42:53,2018-01-01 03:42:53,[p]
... if the SX1276 chip inside states to perform -146dBm with a bandwidth of 10.4KHz, while the minimum bandwidth value supported by RN2483 API is 125KHz?

Hi everybody (happy new year!).

According to RN2483 datasheet, its best sensitivity performance in LoRa modulation is -146dBm. However, in order to achieve such a sensitivity, the SX1276 transceiver inside the RN2483 module requires a receive bandwidth of 10.4KHz or less.
Unfortunately, the RN2483 API does only support bw configuration of 500, 250 and 125 KHz. Using a bw value of 125KHz the SX1276 transceiver should be capable to perform a maximum sensitivity of -136dBm.

Am I missing something, or is there actually some trouble with RN2483 datasheet (or command interface) ?

[/p],0,0,0,0

当前读取代码

with open(pathCsv, encoding="utf-16") as f:
    csv_reader = csv.DictReader(f)
    for row in csv_reader :
        print(row.get('commentText'))
        break

错误输出

[p]

解决方案

核心思路是识别commentText字段的起止标记([p]和[/p]),合并分散的多行内容并为字段添加双引号,让CSV解析器能正确识别完整字段。以下是两种可行方法:

方法1:预处理文件修复格式

生成一个修复后的CSV文件,后续直接用标准工具读取:

def fix_corrupted_csv(input_path, output_path, encoding="utf-16"):
    with open(input_path, 'r', encoding=encoding) as infile, open(output_path, 'w', encoding=encoding) as outfile:
        # 写入表头
        header = infile.readline().strip()
        outfile.write(f"{header}\n")
        
        current_row = []
        in_comment = False
        comment_content = []
        
        for line in infile:
            stripped_line = line.strip()
            if not stripped_line:
                if in_comment:
                    comment_content.append(line)
                continue
            
            # 进入commentText字段
            if '[p]' in stripped_line and not in_comment:
                parts = stripped_line.split(',', 24)  # 分割到commentText字段开头
                current_row = parts[:24]
                comment_content.append(parts[24])
                in_comment = True
            # 结束commentText字段
            elif '[/p]' in stripped_line and in_comment:
                comment_content.append(stripped_line.split('[/p]')[0] + '[/p]')
                # 拼接完整行,给commentText加双引号
                fixed_comment = '"' + ''.join(comment_content).replace('"', '""') + '"'
                remaining_parts = stripped_line.split('[/p]')[1].split(',')
                full_row = ','.join(current_row) + ',' + fixed_comment + ',' + ','.join(remaining_parts)
                outfile.write(f"{full_row}\n")
                # 重置状态
                current_row = []
                comment_content = []
                in_comment = False
            # 收集comment中间内容
            elif in_comment:
                comment_content.append(stripped_line + '\n')

# 使用示例
fix_corrupted_csv("corrupted.csv", "fixed.csv")

# 读取修复后的文件
with open("fixed.csv", encoding="utf-16") as f:
    reader = csv.DictReader(f)
    for row in reader:
        print(row["commentText"])
        break

方法2:自定义解析逻辑直接读取

无需生成新文件,在读取过程中实时合并字段内容:

import csv

def read_corrupted_csv(path, encoding="utf-16"):
    rows = []
    with open(path, 'r', encoding=encoding) as f:
        header = f.readline().strip().split(',')
        comment_col_idx = header.index("commentText")
        current_row = {}
        in_comment = False
        comment_buffer = []
        
        for line in f:
            stripped_line = line.strip()
            if not in_comment:
                parts = line.strip().split(',', comment_col_idx + 1)
                if len(parts) > comment_col_idx and '[p]' in parts[comment_col_idx]:
                    # 初始化当前行的前半部分
                    for idx in range(comment_col_idx):
                        current_row[header[idx]] = parts[idx]
                    comment_buffer.append(parts[comment_col_idx])
                    in_comment = True
            else:
                if '[/p]' in stripped_line:
                    # 结束comment字段,拼接内容
                    comment_buffer.append(stripped_line.split('[/p]')[0] + '[/p]')
                    current_row["commentText"] = ''.join(comment_buffer)
                    # 处理剩余字段
                    remaining = stripped_line.split('[/p]')[1].split(',')
                    for idx in range(comment_col_idx + 1, len(header)):
                        pos = idx - (comment_col_idx + 1)
                        current_row[header[idx]] = remaining[pos] if pos < len(remaining) else ""
                    rows.append(current_row)
                    # 重置状态
                    current_row = {}
                    comment_buffer = []
                    in_comment = False
                else:
                    comment_buffer.append(line.rstrip('\n'))
    return rows

# 使用示例
data = read_corrupted_csv("corrupted.csv")
print(data[0]["commentText"])

内容的提问来源于stack exchange,提问作者Van Jake

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.07 23:05:25