Pandas:将分组结果转为DataFrame并合并相似索引重复行
Hey there! Let's sort out this duplicate merging task for your DataFrame. I see you want to combine rows that share the same similarity_index—here's how to do it properly:
First, let's recap your original DataFrame for clarity:
import pandas as pd data = { "ad soyad": ["hakan özdemir", "hasan yaman", "naci şenli", "naciye şen", "osman uygur", "elif sözen", "irem derici"], "similarity_index": [0, 1, 2, 2, 3, 4, 5] } test_df = pd.DataFrame(data)
The code you tried (test_df.set_index("similarity_index").sort_index()) just reorders and indexes your data, but it doesn't actually merge the duplicate rows. To combine rows with matching similarity_index values, we'll use Pandas' groupby method, which is built exactly for this kind of grouping/aggregation task.
Solution 1: Merge names into a single string
If you want to combine the ad soyad entries into a single comma-separated string for each index, use this:
# Group by similarity_index, join the names, and convert back to a DataFrame merged_df = test_df.groupby("similarity_index")["ad soyad"].agg(', '.join).reset_index()
This will give you a clean DataFrame where each unique similarity_index has one row with all corresponding names merged:
similarity_index ad soyad 0 0 hakan özdemir 1 1 hasan yaman 2 2 naci şenli, naciye şen 3 3 osman uygur 4 4 elif sözen 5 5 irem derici
Solution 2: Keep grouped values as lists (for more flexibility)
If you prefer to keep each original entry as an element in a list (useful if you need to process them later), use agg(list) instead:
merged_df = test_df.groupby("similarity_index").agg(list).reset_index()
The output will look like this:
similarity_index ad soyad 0 0 [hakan özdemir] 1 1 [hasan yaman] 2 2 [naci şenli, naciye şen] 3 3 [osman uygur] 4 4 [elif sözen] 5 5 [irem derici]
The key here is using groupby to cluster rows by similarity_index, then applying an aggregation function to combine the values in each group. The reset_index() step ensures your result is a proper DataFrame (not a Series with the index as similarity_index).
内容的提问来源于stack exchange,提问作者Hakan Utku Özdemir

