You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何高效将Pandas DataFrame中的问卷文本答案替换为数值?

高效替换Pandas问卷文本答案为数值的方法

针对你的1800条问卷数据,以下是几种高效的文本转数值实现方式,均可直接处理object类型的列:

1. 字典映射 + map()(单列/指定列首选)

这是最直观且性能优异的方式,适合单独处理某几列:

# 定义映射字典
answer_mapping = {'agree': 1, 'neutral': 2, 'disagree': 3}

# 单列替换
df['question_1'] = df['question_1'].map(answer_mapping)

# 多列批量替换(比如所有以'question_'开头的列)
question_cols = [col for col in df.columns if col.startswith('question_')]
for col in question_cols:
    df[col] = df[col].map(answer_mapping)

map()在处理Series时内部做了优化,1800行数据的处理速度可以忽略不计。

2. replace()方法(多列批量处理更简洁)

如果需要一次性替换DataFrame中所有符合映射规则的列,replace()可以直接传入字典批量操作:

answer_mapping = {'agree': 1, 'neutral': 2, 'disagree': 3}

# 替换所有包含目标文本的列
df.replace(answer_mapping, inplace=True)

# 或者指定列替换
df[question_cols] = df[question_cols].replace(answer_mapping)

replace()底层是向量化操作,性能和map()接近,代码更简洁。

3. 分类类型(Categorical)转换(适合后续统计分析)

如果问卷答案有明确顺序(比如agree < neutral < disagree),可以将列转为Categorical类型,同时指定顺序,后续统计时更合理:

from pandas.api.types import CategoricalDtype

# 定义分类顺序
cat_type = CategoricalDtype(categories=['agree', 'neutral', 'disagree'], ordered=True)
# 转换为分类类型
df[question_cols] = df[question_cols].astype(cat_type)
# 将分类转为对应数值(默认从0开始,这里+1让agree对应1)
df[question_cols] = df[question_cols].cat.codes + 1

这种方式不仅完成数值替换,还保留了答案的顺序关系,后续做排序、分位数计算等统计操作时更准确,性能完全满足1800行的规模。

注意事项

  • 若存在映射字典外的文本值,map()会返回NaN,replace()则保留原文本;如需处理可添加fillna(),比如df['question_1'].map(answer_mapping).fillna(0)(0代表无效值)。
  • 1800行属于小型数据集,上述三种方法的性能差异几乎可以忽略,可根据后续需求选择最适合的方式。

内容的提问来源于stack exchange,提问作者LM1995

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.26 09:18:30