You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

技术问询:基于ParticipantID匹配值在长格式数据框中新增人口统计信息列

解决方案:保留长格式的同时添加人口统计字段

这问题我之前也碰到过,完全懂你不想把整个表转成宽格式的需求!核心思路是单独提取人口统计信息并转成宽格式(仅这部分),再通过ParticipantID关联回原数据,这样既保留了Q1/Q2这类问题的长格式结构,又能把人口统计值加到对应参与者的每一行里。

下面给你两种常用数据分析工具的实现方法:

R语言(dplyr + tidyr)

假设你的数据框名为df,步骤如下:

  1. 先筛选出人口统计类的问题(比如Age、Gender、Education),把这部分转成宽格式(每个ParticipantID一行,人口统计字段作为列)
  2. 将原数据和这个宽格式的人口统计表按ParticipantID合并
library(dplyr)
library(tidyr)

# 提取人口统计信息并转宽
demographics <- df %>%
  filter(Question %in% c("Age", "Gender", "Education")) %>%
  pivot_wider(names_from = Question, values_from = Resp)

# 合并回原数据,保留原长格式
df_with_demographics <- df %>%
  left_join(demographics, by = "ParticipantID")

Python(pandas)

同样的思路,用pandas实现:

import pandas as pd

# 提取人口统计行,转成宽格式(索引为ParticipantID)
demographics = df[df['Question'].isin(['Age', 'Gender', 'Education'])].pivot(
    index='ParticipantID',
    columns='Question',
    values='Resp'
).reset_index()

# 合并到原数据,保留长格式结构
df_with_demographics = pd.merge(df, demographics, on='ParticipantID', how='left')

这样处理后,你就能得到想要的结果:每个参与者的所有行(包括Q1、Q2等问题的回答行)都会带上对应的Age、Gender、Education等值,同时原数据的长格式完全保留。

内容的提问来源于stack exchange,提问作者TSL-Aquilus

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.27 15:09:08