You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何去除DataFrame中导入文本附带的换行符\n

去除DataFrame合并后Ground_truth字段的换行符

问题说明

合并预测文本DataFrame与真实标签DataFrame后,生成的Class_B.jsonl中Ground_truth字段带有换行符\n及首尾多余空格,需要清理这些冗余字符。

问题原因

使用readlines()读取txt文件时,会保留每行末尾的换行符\n,同时原txt文件中每行首尾可能存在空格,这些都会被带入生成的DataFrame列中。

解决方案

方法1:读取txt时直接处理

读取每行文本时,用strip()方法去除首尾的空白字符(包括换行符、空格、制表符):

# 替换原读取txt的代码段
file1 = open(f'{working_dir}3rd_col.txt', 'r')
Lines = [line.strip() for line in file1.readlines()]
col_3rd = pd.DataFrame(Lines, columns=['Ground_truth'])

方法2:生成DataFrame后批量清理

如果已经生成col_3rd,可以用Pandas字符串方法批量处理列数据:

col_3rd['Ground_truth'] = col_3rd['Ground_truth'].str.strip()

完整修改代码

import pandas as pd
from google.colab import drive
import json

drive.mount('/content/drive')

def load_jsonl(text_path):
    return pd.read_json(
                         path_or_buf = text_path,
                         lines=True
                        )

working_dir = "/content/drive/MyDrive/Class_B/"

df = load_jsonl(f'{working_dir}labels.jsonl')

# 读取txt并清理每行的空白字符
file1 = open(f'{working_dir}3rd_col.txt', 'r')
Lines = [line.strip() for line in file1.readlines()]
col_3rd = pd.DataFrame(Lines, columns=['Ground_truth'])

result = pd.concat([df, col_3rd ], axis=1)

reddit = result.to_dict(orient= "records")
print(type(reddit) , len(reddit))

with open(f"{working_dir}Class_B.jsonl","w") as f:
    for line in reddit:
        f.write(json.dumps(line,ensure_ascii=False) + "\n")

效果验证

修改后生成的Class_B.jsonl将不再包含换行符和多余空格,符合预期格式:

{"image_name": "1.JPG", "text": "Flattery is words of kindness for a", "Ground_truth": "Flattery is words of kindness for a"}
{"image_name": "2.JPG", "text": "potential favor.", "Ground_truth": "potential favor."}

内容的提问来源于stack exchange,提问作者Mohammed

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.15 08:07:41