You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Pandas移除特定垃圾字符及其前后相邻字符?

处理孟加拉语音译文本中的特殊字符问题

我将孟加拉语音译转换为英文后,解析出现了垃圾字符(্),需要移除该字符及其前后各一个相邻字符。

原DataFrame数据

col1        
utto্tor        
dokkho্shin     
muuns্si    

期望输出

col1            col2
utto্tor        uttor
dokkho্shin     dokkhhin
muuns্si        muuni

注:该特殊字符由Avro parser生成,示例代码如下:

reversed_text = avro.reverse("উত্তর")
print(reversed_text)

# 输出: utto্tor

对应的孟加拉语原文与译文对照:

col0        col1
উত্তর       utto্tor
দক্ষিণ      dokkho্shin
মুন্সী         muuns্si

解决方案

使用Pandas结合正则表达式,匹配并移除特殊字符及其前后各一个字符:

import pandas as pd

# 构造原数据
df = pd.DataFrame({
    'col1': ['utto্tor', 'dokkho্shin', 'muuns্si']
})

# 利用正则替换:匹配任意字符+特殊字符(Unicode U+09CD)+任意字符,替换为空
df['col2'] = df['col1'].str.replace(r'.\x{09CD}.', '', regex=True)

# 查看结果
print(df)

执行后输出结果:

col1       col2
0    utto্tor      uttor
1  dokkho্shin  dokkhhin
2    muuns্si      muuni

内容的提问来源于stack exchange,提问作者asif abdullah

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.14 09:40:31