Python中基于列比较DataFrame连续行及供应商数据校验
Hey there! Let's simplify this a lot—Pandas has built-in tools that make checking consistency across Supplier/Grp groups way easier than looping through rows manually.
First, let's ditch that complex loop you wrote for filtering a single supplier. You can slice your DataFrame directly like this:
su = "Authentic Brands Group LLC" # Filter for your target supplier and keep only the columns you need df3 = data[data['Supplier'] == su][['Supplier','ID','Units','Grp']].copy()
That's way cleaner and faster than building the DataFrame row-by-row.
Now, to add a column that marks whether Units are consistent for each (Supplier, Grp) pair, we can use groupby() and transform() to check how many unique Units values exist in each group. If the count is 1, all values are consistent; if it's more than 1, they're inconsistent.
Here's the code to add the new column (let's call it Units_Consistent):
# Add a consistency flag column df3['Units_Consistent'] = df3.groupby(['Supplier', 'Grp'])['Units'].transform( lambda x: x.nunique() == 1 )
Let's break this down for you:
groupby(['Supplier', 'Grp']): Groups the DataFrame by every unique combination of Supplier and Grp.['Units'].transform(...): Runs the lambda function on each group's Units column, and returns a value for every row (so it matches the original DataFrame's row count).lambda x: x.nunique() == 1: Checks if there's only 1 unique value in the group's Units. ReturnsTrueif all values match,Falseif there are inconsistencies.
Example with sample data:
If your df3 looks like this:
| Supplier | ID | Units | Grp |
|---|---|---|---|
| Authentic Brands Group LLC | 1 | 10 | GrpA |
| Authentic Brands Group LLC | 2 | 10 | GrpA |
| Authentic Brands Group LLC | 3 | 20 | GrpA |
| Authentic Brands Group LLC | 4 | 15 | GrpB |
| Authentic Brands Group LLC | 5 | 15 | GrpB |
The resulting Units_Consistent column will be:
| Units_Consistent |
|---|
| False |
| False |
| False |
| True |
| True |
Bonus: Check all suppliers at once
If you ever want to check consistency across all suppliers (not just one), you can run this on your original data DataFrame directly—no need to split into individual supplier DataFrames:
data['Units_Consistent'] = data.groupby(['Supplier', 'Grp'])['Units'].transform( lambda x: x.nunique() == 1 )
This will add the consistency flag to every row in your original dataset.
内容的提问来源于stack exchange,提问作者frisbeee

