如何在DataFrame字符串列中精准匹配含空格词组并生成标记列
Got it, let's tackle this problem head-on. The core issue here is matching entire space-containing phrases (not just individual words) from your list against the text column—splitting the text into words won't work because we need to keep those space-separated phrases intact as single match targets.
Step-by-Step Implementation
Let's walk through this using your sample data and requirements:
Set up your sample data and target phrase list
First, let's recreate your DataFrame and the provided phrase list:import pandas as pd # Sample DataFrame data = { "id": ["a", "b", "c", "d", "e", "f"], "text": [ "simultaneous there the", "simultaneous there", "mul why the", "mul the", "simul a b c", "a c b" ] } df = pd.DataFrame(data) # Phrases we need to match (with spaces) list_provided = ["mul the", "a b c"]Build a regex pattern for exact phrase matching
We'll combine all phrases in the list into a single regex pattern using|(which means "or" in regex). This lets us check if any of the full phrases exist as a substring in the text column:# Join phrases into a regex "or" pattern match_pattern = "|".join(list_provided)Add the
foundcolumn
Usestr.contains()to detect matches, then convert the boolean results to integers (1 for matches, 0 for non-matches):df["found"] = df["text"].str.contains(match_pattern).astype(int)
Final Result
Running this code will produce your expected output:
| id | text | found |
|---|---|---|
| a | simultaneous there the | 0 |
| b | simultaneous there | 0 |
| c | mul why the | 0 |
| d | mul the | 1 |
| e | simul a b c | 1 |
| f | a c b | 0 |
Handling Special Characters (Optional)
If your phrase list includes regex special characters (like ., *, or ?), you'll need to escape them to avoid unintended matches. Use re.escape() to sanitize each phrase:
import re # Escape special characters in phrases before building the pattern match_pattern = "|".join(re.escape(phrase) for phrase in list_provided) df["found"] = df["text"].str.contains(match_pattern).astype(int)
This ensures phrases with special characters are matched exactly as written, not interpreted as regex syntax.
内容的提问来源于stack exchange,提问作者Sam

