You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于ID替换Pandas DataFrame中每组首行指定文本的方法

按ID分组替换首次出现的"review"

问题分析

你之前的代码错误在于np.argmax(df.values=="review",1)会遍历每一行查找第一个匹配"review"的列,导致所有行的对应位置都被替换,而非每组仅替换第一行。

可行解法

方法一:用分组累计计数标记首行

import pandas as pd

# 构造你的数据集
data = {
    'ID': [1895001, 1895001, 1895001, 2104264, 2102404, 2102404, 1809905, 1809905, 1809905, 1811700],
    'Status': ['review'] * 10
}
df = pd.DataFrame(data)

# 给每组的第一行打标记
df['is_first'] = df.groupby('ID').cumcount() == 0

# 只替换每组首行且Status为review的记录
df.loc[df['is_first'] & (df['Status'] == 'review'), 'Status'] = 'first review'

# 删掉辅助列(可选)
df.drop('is_first', axis=1, inplace=True)

print(df)

方法二:分组后直接修改每组首条匹配记录

import pandas as pd

data = {
    'ID': [1895001, 1895001, 1895001, 2104264, 2102404, 2102404, 1809905, 1809905, 1809905, 1811700],
    'Status': ['review'] * 10
}
df = pd.DataFrame(data)

def update_first_review(group):
    # 找到组内第一个Status为review的索引
    first_match_idx = group[group['Status'] == 'review'].index[0]
    group.loc[first_match_idx, 'Status'] = 'first review'
    return group

# 分组应用修改函数
df = df.groupby('ID', group_keys=False).apply(update_first_review)

print(df)

输出结果

两种方法都会得到你期望的数据集:

ID         Status
0  1895001  first review
1  1895001        review
2  1895001        review
3  2104264  first review
4  2102404  first review
5  2102404        review
6  1809905  first review
7  1809905        review
8  1809905        review
9  1811700  first review

内容的提问来源于stack exchange,提问作者dagi_de

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.31 14:41:43