You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python新手求助:特征工程中姓名词数判断特征列的代码实现

解决Name字段词数判断特征列Feat_7的语法问题

需求说明

需要构建特征列Feat_7:当Name字段的词数小于2或大于2时,值为True;词数等于2时为False。

示例数据

NameStreetCity
UAB 3 berzeliaiTestine g. 1Vilnius
UAB ACCTerminalo g. 8Biruliskes, Kauno Raj
ACCTerminalo g. 8Biruliskes, Kauno Raj
ACMETerminalo g. 8Biruliskes, Kauno Raj
IĮ RododendrasTestine g. 9Biruliskes, Kauno Raj

原代码问题点

原Feat_7的代码存在两处语法错误:

  1. 错误调用['Name'.split()],应该对当前遍历的x(即Name字段的单条值)执行拆分,而非字符串'Name'
  2. Lambda表达式中不能直接写if语句,需用布尔表达式或条件表达式

原错误代码片段:

df['Feat_7'] = df['Name'].apply(lambda x: **if len(['Name'.split()]) <2** <--???? in x)

修正后的完整代码

import pandas as pd

# 假设df是你的数据集
df['Feat_1'] = df['Name'].apply(lambda x: "UAB" in x)
df['Feat_2'] = df['Name'].apply(lambda x: "AB" in x)
df['Feat_3'] = df['Name'].apply(lambda x: "MB" in x)
df['Feat_4'] = df['Name'].apply(lambda x: "II" in x)
df['Feat_5'] = df['Name'].apply(lambda x: "IĮ" in x)
df['Feat_6'] = df['Name'].apply(lambda x: "SB" in x)

# 正确的Feat_7写法:判断词数不等于2
df['Feat_7'] = df['Name'].apply(lambda x: len(x.split()) != 2)

print(df.head())

代码解释

  • x.split():将单条Name值按空格拆分成词列表
  • len(x.split()):获取词的数量
  • != 2:直接判断词数是否不等于2,满足则返回True,否则False,完全匹配需求

对应示例数据的Feat_7结果

NameFeat_7
UAB 3 berzeliaiTrue
UAB ACCFalse
ACCTrue
ACMETrue
IĮ RododendrasFalse

内容的提问来源于stack exchange,提问作者Paulius Banys

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.06 07:35:11