如何用Pandas从多语言字符串中提取拉丁词至单独列?
问题解决:从多语言字符串提取拉丁词到单独列
错误原因
你直接用re.findall()处理Pandas的Series对象(DataFrame的列),但re.findall()仅接受单个字符串/字节类型的输入,因此触发TypeError: expected string or bytes-like object错误。
正确实现代码
要提取所有拉丁词并生成新列,可借助Pandas的字符串方法str.findall()结合正则表达式,匹配连续的拉丁字母(含大小写)、反斜杠、斜杠,再将匹配结果拼接为字符串:
import pandas as pd import re st = {'string':['hhhh 15-0850tcx cord\\with plastic end / light mustard -82cm шнур нужд вес 07 кг','1. 06900000027899 non woven 12 grid socks']} s = pd.DataFrame(st) # 定义正则:匹配连续的拉丁字母、反斜杠、斜杠 pattern = r'[a-zA-Z\\/]+' # 提取每个字符串中的匹配项,并用空格拼接 s['latin_words'] = s['string'].str.findall(pattern).str.join(' ') # 输出合并后的目标字符串 print(' '.join(s['latin_words'].tolist()))
输出结果
hhhh tcx cord\with plastic end / light mustard cm non woven grid socks
说明
- 使用
str.findall()而非原生re.findall(),可直接对Series的每个元素批量处理 - 正则
[a-zA-Z\\/]+匹配连续的拉丁字母、反斜杠、斜杠,过滤掉数字、西里尔字母等无关内容 - 用
str.join(' ')将提取到的多个片段拼接成完整字符串,最后合并两列结果得到目标输出
内容的提问来源于stack exchange,提问作者Madina
相关产品推荐
相关产品推荐

