如何检查DataFrame列是否含多组字符串,生成匹配主题的行记录
Solution to Match Topics with Data Entries
Got it, let's work through this problem. You need to check each entry in df_Data against the keyword lists in df_Topics, then generate a table where every matched topic gets its own row for the corresponding ID. Here's a straightforward way to do it with pandas:
Step-by-Step Code Implementation
First, let's start with your original data (I fixed a small typo in the Red keywords—pinnk → pink to make the example work correctly), then build the solution:
import pandas as pd # Initialize your input DataFrames df_Data = pd.DataFrame({ 'ID': ['123','456','789','100','200'], 'Names': ['the dog is Red and blue','Cat is Pink','animal is cyan','pet is BLUE','i am green'] }) df_Topics = pd.DataFrame({ 'Blue': ['blue','cyan','aqua'], 'Red': ['red','pink','fuscia','crimson'] }) # 1. Add a lowercase version of the Names column to avoid case-sensitive misses df_Data['names_lower'] = df_Data['Names'].str.lower() # 2. Define a helper function to find all matching topics for a single row def find_topics(row): matched_topics = [] # Loop through each topic in df_Topics for topic in df_Topics.columns: # Check if any keyword for this topic exists in the lowercase name keywords = [kw.lower() for kw in df_Topics[topic].dropna()] if any(kw in row['names_lower'] for kw in keywords): matched_topics.append(topic) return matched_topics # 3. Apply the helper function to get all matching topics per ID df_Data['matched_topics'] = df_Data.apply(find_topics, axis=1) # 4. Explode the list of topics into individual rows final_output = df_Data.explode('matched_topics').rename(columns={'matched_topics': 'Topics'}) # 5. Clean up to keep only the required columns final_output = final_output[['ID', 'Topics']].dropna().reset_index(drop=True) # Print the result print(final_output)
Sample Output
Running this code will give you exactly the table you're looking for:
ID Topics 0 123 Blue 1 123 Red 2 456 Red 3 789 Blue 4 100 Blue
Key Notes
- Case Insensitivity: We convert both the
Namestext and keywords to lowercase to ensure matches regardless of capitalization (like matchingBLUEfromdf_Datatoblueindf_Topics). - Handling Missing Keywords: The
.dropna()call ensures we skip any empty values indf_Topicsthat might cause false matches. - Explode for Row Expansion: The
explode()method is perfect here—it takes each topic in the list for an ID and turns it into a separate row, which is exactly the format you need.
内容的提问来源于stack exchange,提问作者jvirgi
相关产品推荐
相关产品推荐

