如何用Python提取两列表匹配词及处理DataFrame列内列表匹配生成新列
Hey there! Let's tackle your two Python challenges step by step—first the basic list comparison, then the trickier DataFrame problem you're stuck on.
The simplest way to find common words between two lists is using sets, since they’re optimized for membership checks and intersections. Here's how it works:
list_a = ["apple", "banana", "cherry", "date"] list_b = ["banana", "date", "fig", "grape"] # Get matching words (order might not be preserved) matches = list(set(list_a) & set(list_b)) print(matches) # Output: ['banana', 'date'] (order varies because sets are unordered)
If you need to keep the order from one of the original lists, use a list comprehension instead:
# Preserve order from list_a ordered_matches = [word for word in list_a if word in list_b] print(ordered_matches) # Output: ['banana', 'date']
First, let's recreate your sample DataFrame to test our solution:
import pandas as pd # Sample data matching your example data = { 'ticket_subject': [["mouse", "is not", "working"], ["password", "wrong"], ["reset", "password"], ["forgot", "userid"]], 'keys': [["hardware", "change"], ["error", "password"], ["account", "lock"], ["userid", "forgot"]], 'automation': ["70%", "50%", "50%", "50%"] } df = pd.DataFrame(data)
Solution Based on Your Requirement
Your goal is: For each list in the keys column, collect words that exist in any row of the ticket_subject column. Here's an efficient way to do this:
- First, create a set of all words from the entire
ticket_subjectcolumn (sets make lookups lightning fast):
# Flatten all lists in ticket_subject into a single set of words all_subject_words = set(word for sublist in df['ticket_subject'] for word in sublist)
- Define a function to filter each
keyslist against this set, then apply it to the column:
def filter_matching_words(key_list): return [word for word in key_list if word in all_subject_words] # Add the new column to your DataFrame df['matching_words'] = df['keys'].apply(filter_matching_words)
If you print the DataFrame now, you’ll get exactly your expected output:
ticket_subject keys automation matching_words 0 [mouse, is not, working] [hardware, change] 70% [] 1 [password, wrong] [error, password] 50% [password] 2 [reset, password] [account, lock] 50% [] 3 [forgot, userid] [userid, forgot] 50% [userid, forgot]
Alternative: Match Words Only in the Same Row
If you ever need to check if words from keys exist only in the same row's ticket_subject list (instead of any row), use this approach instead:
def filter_matching_same_row(row): subject_words = set(row['ticket_subject']) return [word for word in row['keys'] if word in subject_words] df['matching_words_same_row'] = df.apply(filter_matching_same_row, axis=1)
This gives the same result for your sample data, but it’s a handy option if your requirements shift later.
内容的提问来源于stack exchange,提问作者lakshith

