R中grepl结合dplyr filter的Python等价实现技术问询
Got it, let's sort this out! The problem with using isin() here is that it only returns rows where the entire Countries string exactly matches one of your target values. But we need to catch rows where either Brazil or Argentina appears anywhere in the comma-separated list—just like your R code does with grepl or the strsplit/lapply approach.
Here are two pandas solutions that mirror your R implementations:
Solution 1: Using str.contains() (Equivalent to dplyr + grepl)
This uses regex to search for either "Brazil" or "Argentina" as a substring in the Countries column, just like grepl('(Brazil|Argentina)', Countries) in R:
import pandas as pd df = pd.DataFrame(dict( Countries=['Brazil','Venezuela','Brazil, Colombia, Paraguay','Argentina','Peru','Andorra,Argentina,Chile,Uruguay'], Code=[1,2,3,4,5,6] )) # Filter rows where Countries contains Brazil or Argentina filtered_df = df[df['Countries'].str.contains('Brazil|Argentina', regex=True)] print(filtered_df)
Solution 2: Split, Clean, and Check Membership (Mirroring strsplit/lapply)
This replicates your second R approach: splitting the string, cleaning whitespace, then checking if any element matches your target countries:
import pandas as pd df = pd.DataFrame(dict( Countries=['Brazil','Venezuela','Brazil, Colombia, Paraguay','Argentina','Peru','Andorra,Argentina,Chile,Uruguay'], Code=[1,2,3,4,5,6] )) targets = {'Brazil', 'Argentina'} # Using a set for faster membership checks # For each row: split by commas, strip spaces, check if any match targets has_target = df['Countries'].apply( lambda row: any(country.strip() in targets for country in row.split(',')) ) filtered_df = df[has_target] print(filtered_df)
Expected Output for Both Solutions
Countries Code 0 Brazil 1 2 Brazil, Colombia, Paraguay 3 3 Argentina 4 5 Andorra,Argentina,Chile,Uruguay 6
Both methods will give you the exact result you're looking for, matching the output from your R code.
内容的提问来源于stack exchange,提问作者bzimons

