在Pandas中按用户对分组提取指定行数数据
Got it, let's solve this problem efficiently. Your DataFrame is structured such that every 6 rows represent a pair of users sharing movie ratings, and you need to grab the first 4 rows from each of these groups. Here's how to do it:
Step 1: Understand the Grouping Logic
Each user pair's data occupies exactly 6 consecutive rows. So we can group the DataFrame using integer division of the index by 6—this will bundle rows 0-5 into group 0, 6-11 into group 1, and so on, which matches your data's grouping pattern.
Step 2: Implement the Solution
Use Pandas' groupby() combined with head(4) to extract the first 4 rows from each group:
from pandas import DataFrame import pandas as pd # Original DataFrame df = pd.DataFrame({ 'userId': [10,20,10,20,10,20,60,90,60,90,60,90,30,40,30,40,30,40,50,60,50,60,50,60], 'movieId': [500,500,800,800,700,700,1100,1100,1900,1900,2000,2000,1600,1600,1901,1901,3000,3000,3025,3025,4000,4000,500,500], 'ratings': [3.5,4.5,2.0,5.0,4.0,1.5,3.5,4.5,3.5,4.5,2.0,5.0,4.0,1.5,3.5,4.5,3.5,4.5,2.0,5.0,4.0,1.5,3.5,4.5] }) # Group by every 6 rows, take first 4 rows from each group result = df.groupby(df.index // 6).head(4) # Print the result to verify print(result)
Step 3: Verify the Output
Running this code will give you exactly the expected result:
userId movieId ratings 0 10 500 3.5 1 20 500 4.5 2 10 800 2.0 3 20 800 5.0 6 60 1100 3.5 7 90 1100 4.5 8 60 1900 3.5 9 90 1900 4.5 12 30 1600 4.0 13 40 1600 1.5 14 30 1901 3.5 15 40 1901 4.5 18 50 3025 2.0 19 60 3025 5.0 20 50 4000 4.0 21 60 4000 1.5
This approach works perfectly even when users appear in multiple groups (like userId=60) because we're strictly grouping by row positions, which aligns with your data's predefined structure.
内容的提问来源于stack exchange,提问作者Sanwal

