You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何检查DataFrame列是否含多组字符串,生成匹配主题的行记录

Solution to Match Topics with Data Entries

Got it, let's work through this problem. You need to check each entry in df_Data against the keyword lists in df_Topics, then generate a table where every matched topic gets its own row for the corresponding ID. Here's a straightforward way to do it with pandas:

Step-by-Step Code Implementation

First, let's start with your original data (I fixed a small typo in the Red keywords—pinnk → pink to make the example work correctly), then build the solution:

import pandas as pd

# Initialize your input DataFrames
df_Data = pd.DataFrame({
    'ID': ['123','456','789','100','200'], 
    'Names': ['the dog is Red and blue','Cat is Pink','animal is cyan','pet is BLUE','i am green']
})
df_Topics = pd.DataFrame({
    'Blue': ['blue','cyan','aqua'], 
    'Red': ['red','pink','fuscia','crimson']
})

# 1. Add a lowercase version of the Names column to avoid case-sensitive misses
df_Data['names_lower'] = df_Data['Names'].str.lower()

# 2. Define a helper function to find all matching topics for a single row
def find_topics(row):
    matched_topics = []
    # Loop through each topic in df_Topics
    for topic in df_Topics.columns:
        # Check if any keyword for this topic exists in the lowercase name
        keywords = [kw.lower() for kw in df_Topics[topic].dropna()]
        if any(kw in row['names_lower'] for kw in keywords):
            matched_topics.append(topic)
    return matched_topics

# 3. Apply the helper function to get all matching topics per ID
df_Data['matched_topics'] = df_Data.apply(find_topics, axis=1)

# 4. Explode the list of topics into individual rows
final_output = df_Data.explode('matched_topics').rename(columns={'matched_topics': 'Topics'})

# 5. Clean up to keep only the required columns
final_output = final_output[['ID', 'Topics']].dropna().reset_index(drop=True)

# Print the result
print(final_output)

Sample Output

Running this code will give you exactly the table you're looking for:

ID Topics
0  123   Blue
1  123    Red
2  456    Red
3  789   Blue
4  100   Blue

Key Notes

  • Case Insensitivity: We convert both the Names text and keywords to lowercase to ensure matches regardless of capitalization (like matching BLUE from df_Data to blue in df_Topics).
  • Handling Missing Keywords: The .dropna() call ensures we skip any empty values in df_Topics that might cause false matches.
  • Explode for Row Expansion: The explode() method is perfect here—it takes each topic in the list for an ID and turns it into a separate row, which is exactly the format you need.

内容的提问来源于stack exchange,提问作者jvirgi

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.09 21:42:41