如何在Python 3或Pandas中优化列表/DataFrame列的重复整数(保留重复ID)并生成期望列表?
Got it, let's solve this problem where you need to map duplicate IDs in a list (or DataFrame column) to sequential integers—all while keeping the duplicate occurrences intact. Below are two practical approaches, depending on whether you prefer pure Python or using Pandas:
If you're working with a regular list and don't want to import external libraries, a dictionary-based mapping works perfectly. We'll track each unique ID and assign it a sequential number as we iterate through the list:
original_array = [1, 1, 5, 8, 8, 20213, 22170, 22170] id_mapping = {} current_sequence = 1 result_array = [] for id_val in original_array: # Assign a new sequence number if the ID hasn't been seen before if id_val not in id_mapping: id_mapping[id_val] = current_sequence current_sequence += 1 # Append the mapped number to the result result_array.append(id_mapping[id_val]) print(result_array) # Output: [1, 1, 2, 3, 3, 4, 5, 5]
How it works:
- We use
id_mappingto store each unique original ID and its corresponding sequential number. - Every time we encounter a new ID, we assign it the next available sequence number and increment the counter.
- Duplicate IDs reuse their already assigned number, so duplicates are preserved in the result.
If you're working with a Pandas DataFrame (which is common for tabular data with ID columns), the factorize() method is a built-in, efficient solution. It's designed exactly for this kind of categorical-to-sequential mapping:
import pandas as pd # Sample DataFrame with the original ID column df = pd.DataFrame({ 'original_id': [1, 1, 5, 8, 8, 20213, 22170, 22170] }) # Factorize the column (default starts at 0, so we add 1 to start from 1) df['new_id'] = pd.factorize(df['original_id'])[0] + 1 # Convert the result back to a list if needed print(df['new_id'].tolist()) # Output: [1, 1, 2, 3, 3, 4, 5, 5]
How it works:
pd.factorize()returns two values: an array of integer codes for each element (matching the order of first occurrence), and an array of unique original values.- We take the first returned array and add 1 to shift the sequence from starting at 0 to starting at 1 (remove the
+1if you want 0-based numbering). - This method is optimized for large datasets, making it faster than manual loops for big DataFrames.
内容的提问来源于stack exchange,提问作者Wruktarr

