如何在DataFrame中通过精确匹配指定单词列表创建新列
How to Create a Binary Column for Exact Word Matches in a Pandas DataFrame
Hey there! I get exactly what you're trying to build—you need a new column that flags 1 whenever any word from your target list exactly shows up in the text column, and 0 otherwise. Let's walk through two solid, easy-to-implement solutions:
Approach 1: Using apply() with a Lambda Function
This method is super straightforward and readable, perfect for small to medium-sized datasets:
import pandas as pd # Your sample DataFrame df = pd.DataFrame({ 'id': ['a', 'b', 'c'], 'text': ['simultaneous there the', 'simultaneous there', 'mul why'] }) # Target word list list_provided = ["mul", "the"] # Create the 'found' column df['found'] = df['text'].apply( lambda x: 1 if any(word in list_provided for word in x.split()) else 0 ) # Print the result print(df)
How this works:
x.split()breaks eachtextstring into individual words (splits on whitespace by default)any(word in list_provided for word in x.split())checks if any of those words is an exact match for an entry in your target list- We convert that boolean check to a 1 or 0 using a simple conditional statement
Approach 2: Vectorized Regex with str.contains()
For larger datasets, this vectorized method is way more efficient (no slow row-by-row processing):
import pandas as pd df = pd.DataFrame({ 'id': ['a', 'b', 'c'], 'text': ['simultaneous there the', 'simultaneous there', 'mul why'] }) list_provided = ["mul", "the"] # Build a regex pattern with word boundaries to enforce exact matches pattern = r'\b(' + '|'.join(list_provided) + r')\b' # Create the 'found' column df['found'] = df['text'].str.contains(pattern, regex=True).astype(int) print(df)
How this works:
r'\b'is a regex word boundary—this ensures we only match whole words (so "them" won't incorrectly trigger a match for "the")'|'.join(list_provided)combines your target words into a regex OR pattern (e.g.,mul|the)str.contains()checks each string for the pattern, returns a boolean series, which we convert to integers (True → 1, False → 0)
Expected Output
Both methods will give you exactly the result you're looking for:
id text found 0 a simultaneous there the 1 1 b simultaneous there 0 2 c mul why 1
内容的提问来源于stack exchange,提问作者Sam
相关产品推荐
相关产品推荐

