能否用np.where在DataFrame中逐行实时更新New Sub.列值?
解决方案
你的问题核心在于np.where是向量化批量计算,所有行的计算都依赖列的原始值,无法获取逐行更新后的中间结果,因此没法实现这种依赖前一行状态的累积操作。下面提供两种可行的实现方式:
1. 逐行迭代更新(适合小数据集)
先初始化目标列,再按学生分组后逐行判断并更新累积结果:
import pandas as pd # 示例DataFrame data = { "Student": ["Alice", "Alice", "Alice", "Bob", "Bob"], "Subject.": ["Math", "English", "History", "Math", "Science"], "Mark.": ["pass", "fail", "pass", "pass", "pass"] } Classroom = pd.DataFrame(data) # 初始化New Sub.列 Classroom["New Sub."] = "" # 按学生分组处理累积 for student, group in Classroom.groupby("Student"): current_pass = "" for idx, row in group.iterrows(): if row["Mark."] == "pass": current_pass = f"{current_pass},{row['Subject.']}" if current_pass else row["Subject."] Classroom.loc[idx, "New Sub."] = current_pass
运行后结果符合预期:
| Student | Subject. | Mark. | New Sub. |
|---|---|---|---|
| Alice | Math | pass | Math |
| Alice | English | fail | Math |
| Alice | History | pass | Math,History |
| Bob | Math | pass | Math |
| Bob | Science | pass | Math,Science |
2. 分组+自定义累积函数(适合大数据集)
用groupby结合apply实现更高效的分组累积,避免全局迭代的性能损耗:
def accumulate_pass_subjects(group): accumulated = [] current = "" for _, row in group.iterrows(): if row["Mark."] == "pass": current = f"{current},{row['Subject.']}" if current else row["Subject."] accumulated.append(current) group["New Sub."] = accumulated return group # 分组应用累积函数 Classroom = Classroom.groupby("Student", group_keys=False).apply(accumulate_pass_subjects)
关键说明
- 必须按学生分组处理,否则会把不同学生的科目混在一起累积。
- 两种方法都是逐行维护一个
current_pass变量,实时更新已通过的科目列表,解决了np.where无法获取中间更新值的问题。
内容的提问来源于stack exchange,提问作者MoS
相关产品推荐
相关产品推荐

