You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在DataFrame中通过精确匹配指定单词列表创建新列

How to Create a Binary Column for Exact Word Matches in a Pandas DataFrame

Hey there! I get exactly what you're trying to build—you need a new column that flags 1 whenever any word from your target list exactly shows up in the text column, and 0 otherwise. Let's walk through two solid, easy-to-implement solutions:

Approach 1: Using apply() with a Lambda Function

This method is super straightforward and readable, perfect for small to medium-sized datasets:

import pandas as pd

# Your sample DataFrame
df = pd.DataFrame({
    'id': ['a', 'b', 'c'],
    'text': ['simultaneous there the', 'simultaneous there', 'mul why']
})

# Target word list
list_provided = ["mul", "the"]

# Create the 'found' column
df['found'] = df['text'].apply(
    lambda x: 1 if any(word in list_provided for word in x.split()) else 0
)

# Print the result
print(df)

How this works:

  • x.split() breaks each text string into individual words (splits on whitespace by default)
  • any(word in list_provided for word in x.split()) checks if any of those words is an exact match for an entry in your target list
  • We convert that boolean check to a 1 or 0 using a simple conditional statement

Approach 2: Vectorized Regex with str.contains()

For larger datasets, this vectorized method is way more efficient (no slow row-by-row processing):

import pandas as pd

df = pd.DataFrame({
    'id': ['a', 'b', 'c'],
    'text': ['simultaneous there the', 'simultaneous there', 'mul why']
})

list_provided = ["mul", "the"]

# Build a regex pattern with word boundaries to enforce exact matches
pattern = r'\b(' + '|'.join(list_provided) + r')\b'

# Create the 'found' column
df['found'] = df['text'].str.contains(pattern, regex=True).astype(int)

print(df)

How this works:

  • r'\b' is a regex word boundary—this ensures we only match whole words (so "them" won't incorrectly trigger a match for "the")
  • '|'.join(list_provided) combines your target words into a regex OR pattern (e.g., mul|the)
  • str.contains() checks each string for the pattern, returns a boolean series, which we convert to integers (True → 1, False → 0)

Expected Output

Both methods will give you exactly the result you're looking for:

id                     text  found
0  a  simultaneous there the      1
1  b       simultaneous there      0
2  c                mul why      1

内容的提问来源于stack exchange,提问作者Sam

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.22 09:35:12