You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在DataFrame字符串列中精准匹配含空格词组并生成标记列

Solution: Match Phrases with Spaces and Flag Matches in Pandas

Got it, let's tackle this problem head-on. The core issue here is matching entire space-containing phrases (not just individual words) from your list against the text column—splitting the text into words won't work because we need to keep those space-separated phrases intact as single match targets.

Step-by-Step Implementation

Let's walk through this using your sample data and requirements:

  1. Set up your sample data and target phrase list
    First, let's recreate your DataFrame and the provided phrase list:

    import pandas as pd
    
    # Sample DataFrame
    data = {
        "id": ["a", "b", "c", "d", "e", "f"],
        "text": [
            "simultaneous there the",
            "simultaneous there",
            "mul why the",
            "mul the",
            "simul a b c",
            "a c b"
        ]
    }
    df = pd.DataFrame(data)
    
    # Phrases we need to match (with spaces)
    list_provided = ["mul the", "a b c"]
    
  2. Build a regex pattern for exact phrase matching
    We'll combine all phrases in the list into a single regex pattern using | (which means "or" in regex). This lets us check if any of the full phrases exist as a substring in the text column:

    # Join phrases into a regex "or" pattern
    match_pattern = "|".join(list_provided)
    
  3. Add the found column
    Use str.contains() to detect matches, then convert the boolean results to integers (1 for matches, 0 for non-matches):

    df["found"] = df["text"].str.contains(match_pattern).astype(int)
    

Final Result

Running this code will produce your expected output:

idtextfound
asimultaneous there the0
bsimultaneous there0
cmul why the0
dmul the1
esimul a b c1
fa c b0

Handling Special Characters (Optional)

If your phrase list includes regex special characters (like ., *, or ?), you'll need to escape them to avoid unintended matches. Use re.escape() to sanitize each phrase:

import re

# Escape special characters in phrases before building the pattern
match_pattern = "|".join(re.escape(phrase) for phrase in list_provided)
df["found"] = df["text"].str.contains(match_pattern).astype(int)

This ensures phrases with special characters are matched exactly as written, not interpreted as regex syntax.

内容的提问来源于stack exchange,提问作者Sam

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.22 09:38:51