如何提取pandas DataFrame列中每个元组首元素的第一个单词?
解决方法
你需要先统一所有单词分隔符为空格,再进行分割取首元素,即可同时兼容空格、连字符两种分隔场景:
完整实现代码
import pandas as pd # 原始数据(已修正语法,将两个字典放入列表中) test = [ {'text': [ ('tom-mark', 'tom', 'tom is a good guy.'), ('Nick X','nick', 'Is that Nick?') ]}, {'text': [ ('juli', 'juli', 'Tom likes juli so much.'), ('tony', 'tony', 'Steve and Tony listen in as well.') ]} ] # 合并所有元组构造DataFrame all_records = [] for item in test: all_records.extend(item['text']) df = pd.DataFrame(all_records, columns=['col1', 'col2', 'col3']) # 定义首单词提取函数 def extract_first_word(raw_str): # 将连字符替换为空格,统一分隔规则 processed_str = raw_str.replace('-', ' ') # 按空格分割后取第一个元素 return processed_str.split()[0] # 应用函数提取结果 df['first_word'] = df['col1'].apply(extract_first_word) # 输出结果列表 print(df['first_word'].tolist())
运行后输出结果为:['tom', 'Nick', 'juli', 'tony'],完全符合预期。
扩展说明
如果后续遇到下划线、点号等其他单词分隔符,只需要在替换步骤增加对应的替换逻辑即可,比如同时处理连字符和下划线:
processed_str = raw_str.replace('-', ' ').replace('_', ' ')
如果需要统一输出小写,可以在返回前调用.lower()方法:
return processed_str.split()[0].lower()
内容的提问来源于stack exchange,提问作者Alina
相关产品推荐
相关产品推荐

