You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Pandas正则替换单引号字符串异常:数据合并至首行问题排查

Pandas处理CSV正则替换导致行合并问题排查

问题场景

尝试用Pandas结合正则表达式更新CSV文件,将每行中所有单引号包裹的字符串替换为字面量const,但输出结果中所有数据被合并到第一行,原3行格式无法保留。

原代码

import pandas as pd
import re

df=pd.read_csv("t1.csv");
col1=df['col1']
col2=re.sub(r'\'([^\']*)\'','const',str(col1))
col3 = pd.Series(col2)

df['col1']=col3
df.to_csv('t_u.csv')
exit()

输入t1.csv内容

col1
This one has 'many' 'such' 'quotes' in it.
Now it does not.
But 'this' 'one' does 'have' it 'too'.

错误输出

col1
0   "0    This one has const const const in it.
1                              Now it does not.
2        But const const does const it const.
Name: col1, dtype: object"
1   
2   

期望输出

保持原3行格式,仅替换单引号包裹内容为const:

col1
This one has const const const in it.
Now it does not.
But const const does const it const.

问题原因

核心错误是将整个Series对象转成了字符串(str(col1)),这会把Series的索引、所有行内容以及dtype信息合并成一个大字符串。之后将这个字符串转为Series赋值回原列,只会让第一行填充这个大字符串,其余行因长度不匹配自动补空。

修正方案

不要对整个Series做字符串转换,而是用Series.apply()方法逐行处理每个单元格的字符串,确保每一行的内容都被单独替换:

修正后代码

import pandas as pd
import re

df = pd.read_csv("t1.csv")
# 对col1列的每个元素单独执行正则替换
df['col1'] = df['col1'].apply(lambda x: re.sub(r"'([^']*)'", 'const', x))
df.to_csv('t_u.csv', index=False)  # 添加index=False避免保存时写入索引列

说明

  1. df['col1'].apply(lambda x: ...):遍历col1的每一个单元格,对每个字符串x执行正则替换
  2. index=False:保存CSV时不写入Pandas自动生成的索引列,避免输出出现额外的索引行

内容的提问来源于stack exchange,提问作者Nirav Shah

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.24 22:07:10