You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python中验证Pandas数据框多列行值是否存在于描述列

Pandas验证DataFrame多字段是否存在于描述列并生成差异列

核心需求实现

首先构造示例DataFrame,再通过自定义函数实现核心验证逻辑,生成Validation列:

import pandas as pd

# 构造示例数据
data = {
    'Fruit': ['apple', 'banana', 'grapes', 'pears', 'Strawberry'],
    'Plant': ['Malus domestica', 'Musa', 'Vitis', 'Pyrus calleryana', ''],
    'Location': ['Worldwide', 'Indomalayan realm', 'Northern Hemisphere', 'China', 'Worldwide'],
    'Description': [
        '苹果树在全球广泛种植,是苹果属中栽培最广泛的物种。',
        'Musa属包含可食用香蕉的开花植物,原产于澳大拉西亚。',
        'Vitis属由主要产自北半球的藤本植物物种组成。',
        'Pyrus calleryana是原产于中国的梨树品种。',
        'Fragaria属统称为草莓,在全球广泛种植以获取其果实。'
    ]
}
df = pd.DataFrame(data)

# 定义中英文映射(可根据实际数据调整)
fruit_cn_map = {
    'apple': '苹果',
    'banana': '香蕉',
    'grapes': '葡萄',
    'pears': '梨',
    'Strawberry': '草莓'
}

location_cn_map = {
    'Worldwide': '全球',
    'Indomalayan realm': '澳大拉西亚',
    'Northern Hemisphere': '北半球',
    'China': '中国'
}

# 提取拉丁名属名的辅助函数
def extract_genus(plant_name):
    if not plant_name.strip():
        return ''
    return plant_name.split()[0]

# 核心验证函数
def validate_core(row):
    # 检查水果中文名称是否在描述中
    fruit_ok = fruit_cn_map[row['Fruit']] in row['Description']
    
    # 检查植物拉丁名(或属名)是否在描述中,空值视为通过
    plant = row['Plant'].strip()
    if plant:
        genus = extract_genus(plant)
        plant_ok = (plant in row['Description']) or (genus in row['Description'])
    else:
        plant_ok = True
    
    # 检查地点中文表述是否在描述中
    location_ok = location_cn_map[row['Location']] in row['Description']
    
    # 所有检查项通过则返回True
    return fruit_ok and plant_ok and location_ok

# 生成Validation列
df['Validation'] = df.apply(validate_core, axis=1)

运行后结果与核心期望输出完全匹配。


可选差异列实现

如果需要识别不符项,扩展验证函数同时生成Validation和discrepancy列:

# 带差异识别的验证函数
def validate_with_discrepancy(row):
    discrepancies = []
    
    # 检查水果项
    if fruit_cn_map[row['Fruit']] not in row['Description']:
        discrepancies.append(row['Fruit'])
    
    # 检查植物项:空值标记为缺失,非空则验证是否存在
    plant = row['Plant'].strip()
    if not plant:
        discrepancies.append('Missing Plant')
    else:
        genus = extract_genus(plant)
        if not (plant in row['Description'] or genus in row['Description']):
            discrepancies.append(plant)
    
    # 检查地点项
    if location_cn_map[row['Location']] not in row['Description']:
        discrepancies.append(row['Location'])
    
    # 生成验证结果和差异描述
    validation = len(discrepancies) == 0
    discrepancy_str = ', '.join(discrepancies) if discrepancies else 'N/A'
    
    return pd.Series([validation, discrepancy_str], index=['Validation', 'discrepancy'])

# 应用函数生成两列
df[['Validation', 'discrepancy']] = df.apply(validate_with_discrepancy, axis=1)

运行后会生成包含差异列的结果,与可选输出一致。


内容的提问来源于stack exchange,提问作者Hamza Ahmed

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.28 04:19:58