You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Pandas从多语言字符串中提取拉丁词至单独列?

问题解决:从多语言字符串提取拉丁词到单独列

错误原因

你直接用re.findall()处理Pandas的Series对象(DataFrame的列),但re.findall()仅接受单个字符串/字节类型的输入,因此触发TypeError: expected string or bytes-like object错误。

正确实现代码

要提取所有拉丁词并生成新列,可借助Pandas的字符串方法str.findall()结合正则表达式,匹配连续的拉丁字母(含大小写)、反斜杠、斜杠,再将匹配结果拼接为字符串:

import pandas as pd
import re

st = {'string':['hhhh 15-0850tcx cord\\with plastic end / light mustard -82cm  шнур нужд вес 07 кг','1. 06900000027899 non woven 12 grid socks']}
s = pd.DataFrame(st)

# 定义正则:匹配连续的拉丁字母、反斜杠、斜杠
pattern = r'[a-zA-Z\\/]+'
# 提取每个字符串中的匹配项,并用空格拼接
s['latin_words'] = s['string'].str.findall(pattern).str.join(' ')

# 输出合并后的目标字符串
print(' '.join(s['latin_words'].tolist()))

输出结果

hhhh tcx cord\with plastic end / light mustard cm non woven grid socks

说明

  • 使用str.findall()而非原生re.findall(),可直接对Series的每个元素批量处理
  • 正则[a-zA-Z\\/]+匹配连续的拉丁字母、反斜杠、斜杠,过滤掉数字、西里尔字母等无关内容
  • 用str.join(' ')将提取到的多个片段拼接成完整字符串,最后合并两列结果得到目标输出

内容的提问来源于stack exchange,提问作者Madina

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.10 19:01:17