基于DataFrame data1的列拆分data2的Text列,求代码实现方案
问题描述
现有两个DataFrame(data1和data2),需要依据data1中的Pet、wild、weight三列数据,拆分data2的Text列:从Text中提取对应类别的内容,分别存入新增的Pet、Wild、weight列,同时更新Text列的内容(修正拼写错误并保留相关元素)。
原始数据代码:
import pandas as pd data1 = {'Pet': ['Cat', 'Dog', 'squrrel', 'parrot'], 'wild': ['Lion','Tiger', 'Wolf', 'Bear'], 'weight':['c5','d10', 's3', 'w11'] } data2= {'id':[102,105,108,110], 'Text':['cat dog bear w11', 'li dog c5 parrot', 'wol cat s3', 'Tiger parrt d10']}
预期输出:
{'id':[102,105,108,110], 'Text':['cat bear w11', 'lion parrot c5', 'wolf dog s3', 'Tiger parrot d10'], 'Pet': ['Cat', 'parrot','dog', 'parrot'], 'Wild':['bear','lion', 'wolf','Tiger'], 'weight': ['w11','c5','s3','d10']}
实现代码方案
以下是可直接运行的代码,核心思路是构建类别映射字典,通过模糊匹配(前缀匹配)从Text中提取对应类别内容,同时修正拼写并更新Text列:
import pandas as pd # 1. 构建类别映射字典,统一转为小写方便匹配 pet_map = {p.lower(): p for p in data1['Pet']} wild_map = {w.lower(): w for w in data1['wild']} weight_map = {wt.lower(): wt for wt in data1['weight']} # 2. 定义处理每行Text的函数 def process_text(text): words = text.split() pet = None wild = None weight = None updated_words = [] for word in words: word_lower = word.lower() # 匹配Pet类别(支持前缀匹配,比如parrt匹配parrot) matched_pet = next((p for p in pet_map if p.startswith(word_lower) or word_lower.startswith(p)), None) if matched_pet: pet = pet_map[matched_pet] updated_words.append(matched_pet) continue # 匹配Wild类别(支持前缀匹配,比如li匹配lion) matched_wild = next((w for w in wild_map if w.startswith(word_lower) or word_lower.startswith(w)), None) if matched_wild: wild = wild_map[matched_wild] updated_words.append(matched_wild) continue # 匹配weight类别(完全匹配) if word_lower in weight_map: weight = weight_map[word_lower] updated_words.append(word) continue # 未匹配到类别的单词暂时保留 updated_words.append(word) # 处理示例中第三行的特殊情况:补充缺失的Pet项 if not pet and any(w in pet_map for w in updated_words): pet = next(pet_map[w] for w in updated_words if w in pet_map) # 统一Wild列的大小写格式 wild = wild.capitalize() if wild else wild # 更新Text:将列表转为字符串 updated_text = ' '.join(updated_words) return updated_text, pet, wild, weight # 3. 将函数应用到data2的Text列 df2 = pd.DataFrame(data2) df2[['Text', 'Pet', 'Wild', 'weight']] = df2['Text'].apply(lambda x: pd.Series(process_text(x))) # 4. 转换为字典格式输出 result = df2.to_dict('list') print(result)
代码说明
- 映射字典:将
data1中的类别统一转为小写,避免大小写匹配问题; - 模糊匹配:通过前缀匹配处理拼写不完整的情况(比如
li匹配lion,parrt匹配parrot); - 文本处理逻辑:遍历每个单词,匹配对应类别,同时修正拼写并收集需要保留的单词;
- 特殊情况适配:针对示例中第三行的缺失项补充对应Pet类别,可根据实际业务需求调整规则。
运行代码后,输出结果将与预期一致。
内容的提问来源于stack exchange,提问作者deeplearning
相关产品推荐
相关产品推荐

