You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Python Pandas读取CSV时去除字段外层的=""包裹?

解决Pandas读取带=""包裹字段的CSV问题

针对CSV中部分字段被=""包裹、读取后格式不符合预期的问题,优先在读取阶段处理的方案如下:

方法一:读取时用converters指定列处理函数

自定义清理函数,在读取CSV时直接处理目标列的格式:

import pandas as pd

def clean_quoted_field(s):
    # 仅处理以="开头、以"结尾的字符串
    if isinstance(s, str) and s.startswith('="') and s.endswith('"'):
        return s[2:-1]
    return s

# 读取CSV,指定分隔符为;,并对需要处理的列应用清理函数
df = pd.read_csv('your_file.csv', sep=';', converters={
    'Artist': clean_quoted_field,
    'Title': clean_quoted_field,
    'Price (EUR)': clean_quoted_field,
    'EANCode': clean_quoted_field,
    'InfoLine': clean_quoted_field,
    'PoReference': clean_quoted_field,
    'TotalPrice (EUR)': clean_quoted_field
})

如果不确定哪些列存在格式问题,可以自动对所有列应用函数:

# 先读取表头获取所有列名
cols = pd.read_csv('your_file.csv', sep=';', nrows=0).columns
# 所有列都使用清理函数
df = pd.read_csv('your_file.csv', sep=';', converters={col: clean_quoted_field for col in cols})

方法二:读取后批量处理(备选)

如果读取阶段的处理有局限,可在读取完成后用正则批量清理所有字符串列:

import pandas as pd

df = pd.read_csv('your_file.csv', sep=';')
# 筛选所有字符串类型的列
string_columns = df.select_dtypes(include=['object']).columns
# 用正则替换掉开头的="和结尾的"
df[string_columns] = df[string_columns].apply(lambda x: x.str.replace(r'^="|"$', '', regex=True))

示例验证

针对你提供的CSV:

PackingListNumber;OrderNumber;ArticleId;Artist;Title;Units;MediaType;Price (EUR);EANCode;InfoLine;ReleaseDate;PoReference;ShippedAt;QuantityShipped;TotalPrice (EUR)
2007100976;12669151;1E8085;="WEATHER REPORT";="MR. GONE -COLOURED-";1;LP;="16,9";="8719262030909";="180GR./INSERT/1500 COPIES ON GOLD & BLACK MARBLED VINYL";16-Jun-2023;="";06-Nov-2023;1;="16,9"

处理后,Artist列值为WEATHER REPORT,Price (EUR)为16,9,PoReference变为空字符串,所有字段格式恢复正常。

内容的提问来源于stack exchange,提问作者jose

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.06 07:27:33