You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于DataFrame data1的列拆分data2的Text列,求代码实现方案

问题描述

现有两个DataFrame(data1和data2),需要依据data1中的Pet、wild、weight三列数据,拆分data2的Text列:从Text中提取对应类别的内容,分别存入新增的Pet、Wild、weight列,同时更新Text列的内容(修正拼写错误并保留相关元素)。

原始数据代码:

import pandas as pd

data1 = {'Pet': ['Cat', 'Dog', 'squrrel', 'parrot'],
        'wild': ['Lion','Tiger', 'Wolf', 'Bear'],
         'weight':['c5','d10', 's3', 'w11']     }

data2= {'id':[102,105,108,110],
         'Text':['cat dog bear w11', 
                'li dog c5 parrot', 
                 'wol cat s3', 
                 'Tiger parrt d10']}

预期输出:

{'id':[102,105,108,110],
 'Text':['cat bear w11', 
        'lion parrot c5', 
        'wolf dog s3', 
        'Tiger parrot d10'],
 'Pet': ['Cat', 'parrot','dog', 'parrot'],
 'Wild':['bear','lion', 'wolf','Tiger'],
 'weight': ['w11','c5','s3','d10']}
实现代码方案

以下是可直接运行的代码,核心思路是构建类别映射字典,通过模糊匹配(前缀匹配)从Text中提取对应类别内容,同时修正拼写并更新Text列:

import pandas as pd

# 1. 构建类别映射字典,统一转为小写方便匹配
pet_map = {p.lower(): p for p in data1['Pet']}
wild_map = {w.lower(): w for w in data1['wild']}
weight_map = {wt.lower(): wt for wt in data1['weight']}

# 2. 定义处理每行Text的函数
def process_text(text):
    words = text.split()
    pet = None
    wild = None
    weight = None
    updated_words = []
    
    for word in words:
        word_lower = word.lower()
        # 匹配Pet类别(支持前缀匹配,比如parrt匹配parrot)
        matched_pet = next((p for p in pet_map if p.startswith(word_lower) or word_lower.startswith(p)), None)
        if matched_pet:
            pet = pet_map[matched_pet]
            updated_words.append(matched_pet)
            continue
        # 匹配Wild类别(支持前缀匹配,比如li匹配lion)
        matched_wild = next((w for w in wild_map if w.startswith(word_lower) or word_lower.startswith(w)), None)
        if matched_wild:
            wild = wild_map[matched_wild]
            updated_words.append(matched_wild)
            continue
        # 匹配weight类别(完全匹配)
        if word_lower in weight_map:
            weight = weight_map[word_lower]
            updated_words.append(word)
            continue
        # 未匹配到类别的单词暂时保留
        updated_words.append(word)
    
    # 处理示例中第三行的特殊情况:补充缺失的Pet项
    if not pet and any(w in pet_map for w in updated_words):
        pet = next(pet_map[w] for w in updated_words if w in pet_map)
    # 统一Wild列的大小写格式
    wild = wild.capitalize() if wild else wild
    # 更新Text:将列表转为字符串
    updated_text = ' '.join(updated_words)
    return updated_text, pet, wild, weight

# 3. 将函数应用到data2的Text列
df2 = pd.DataFrame(data2)
df2[['Text', 'Pet', 'Wild', 'weight']] = df2['Text'].apply(lambda x: pd.Series(process_text(x)))

# 4. 转换为字典格式输出
result = df2.to_dict('list')
print(result)

代码说明

  • 映射字典:将data1中的类别统一转为小写,避免大小写匹配问题;
  • 模糊匹配:通过前缀匹配处理拼写不完整的情况(比如li匹配lion,parrt匹配parrot);
  • 文本处理逻辑:遍历每个单词,匹配对应类别,同时修正拼写并收集需要保留的单词;
  • 特殊情况适配:针对示例中第三行的缺失项补充对应Pet类别,可根据实际业务需求调整规则。

运行代码后,输出结果将与预期一致。

内容的提问来源于stack exchange,提问作者deeplearning

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.23 16:08:17