You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Pandas DataFrame中按位置组统计Sample列重复项并添加序号

Solution for Generating Grouped Repeat Counts in Pandas

Perfect, let's tackle this problem step by step. Here's exactly how you can generate that Repeat Number column as you described:

Full Working Code

import pandas as pd

# Your original data
data = {'Sample':['S1', 'S1', 'S1' ,'S1' ,'S2' ,'S2' ,'S3' ,'S3', 'S4', 'Negative', 'Positive', 'Negative', 'S1', 'S1', 'S1' ,'S2' ,'S2' ,'S2' ,'S3' ,'S4', 'S4', 'Positive', 'Positive', 'Negative'], 
        'Location':['A1', 'A2', 'A3' ,'A4' ,'A5' ,'A6' ,'A7' ,'A8', 'A9', 'A10', 'A11', 'A12', 'B1', 'B2', 'B3' ,'B4' ,'B5' ,'B6' ,'B7' ,'B8', 'B9', 'B10', 'B11', 'B12']}
df1 = pd.DataFrame(data)

# Step 1: Create a group identifier (A/B) from the Location column
df1['Group'] = df1['Location'].str[0]

# Step 2: Generate running counts for each Sample within its group
df1['Repeat Number'] = df1.groupby(['Group', 'Sample']).cumcount() + 1

# Step 3: Convert count to string to match your desired output format
df1['Repeat Number'] = df1['Repeat Number'].astype(str)

# Optional: Remove the temporary Group column if not needed
df1 = df1.drop('Group', axis=1)

# View the result
print(df1)

Breakdown of Each Step

  • Create Group Identifier: We extract the first character from Location (e.g., 'A' from 'A1', 'B' from 'B1') to split the data into your A and B groups. This acts as our top-level grouping key.
  • Generate Repeat Counts: Using groupby(['Group', 'Sample']), we first group by the A/B group, then by each unique Sample value. The cumcount() method generates a zero-indexed running count for each subgroup—adding 1 shifts it to start at 1, matching your example.
  • Format to String: Convert the numeric count to a string type to exactly match the output you provided. If numeric values are acceptable for your use case, you can skip this step.
  • Cleanup: Drop the temporary Group column if you don't need to keep it in your final DataFrame.

This will produce exactly the output you shared, with Repeat Number values incrementing correctly within each A/B group for every duplicate Sample entry.

内容的提问来源于stack exchange,提问作者MNVLEY

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.08 22:42:57