Pandas技术问题:如何判断某列值是否存在于另一列的列表中
Hey there! I totally get why isin() isn't working for you here — that method checks if values exist in a global collection (like a single list or another entire column), but you need to do a row-wise check to see if the Country value is inside the specific list in the same row's Another Country's Neighbors column. Let's walk through the right ways to solve this.
First, let's set up an example DataFrame to work with
import pandas as pd # Sample data matching your use case data = { "Country": ["France", "Germany", "Spain", "Italy"], "Another Country's Neighbors": [["Germany", "Spain"], ["France", "Poland"], ["France", "Portugal"], ["Switzerland", "Austria"]] } df = pd.DataFrame(data)
Method 1: Simple row-wise check with apply()
This is the most straightforward approach for small to medium datasets:
# Add the boolean column by checking each row individually df["Country Neighbor"] = df.apply( lambda row: row["Country"] in row["Another Country's Neighbors"], axis=1 # Important! This tells pandas to process row-by-row, not column-by-column )
After running this, your DataFrame will have a new Country Neighbor column with True where the country is in its own row's neighbor list, and False otherwise.
Method 2: Efficient alternative for large datasets
If you're working with a huge DataFrame, apply() can be slow. Instead, use explode() and grouping to avoid row-wise loops:
# Step 1: Expand each neighbor list into individual rows, keeping track of original row indices exploded_neighbors = df.explode("Another Country's Neighbors").reset_index() # Step 2: Mark rows where the country matches a neighbor exploded_neighbors["is_match"] = exploded_neighbors["Country"] == exploded_neighbors["Another Country's Neighbors"] # Step 3: Group by original row index and check if any match exists for that row match_results = exploded_neighbors.groupby("index")["is_match"].any() # Step 4: Attach the results back to the original DataFrame df["Country Neighbor"] = match_results.values
This method is vectorized under the hood, so it's way faster for big datasets.
Why your original isin() attempt failed
When you tried df["Country"].isin(df["Another Country's Neighbors"]), pandas was trying to combine all the neighbor lists into one big collection, then checking if each Country exists anywhere in that collection. That's not what you want — you need to check each country against its own row's list, not every neighbor list in the entire column. Plus, isin() doesn't handle columns filled with lists natively, which is probably why you got an error.
内容的提问来源于stack exchange,提问作者tjones

