如何在Pandas DataFrame中根据指定列的关键词筛选并返回另一列的对应值
Hey there! Let's walk through how to tackle both of your pandas needs clearly and simply.
1. Basic: Extract corresponding names for a single keyword
You already got the first part right with str.contains to create a boolean mask—now we just need to use that mask to slice the DataFrame and pull out the names column.
First, let's set up your sample DataFrame:
import pandas as pd df = pd.DataFrame({ 'names': ['a', 'b', 'c'], 'words': ['apple', 'apple', 'pear'] })
Then, apply your mask and extract the values:
# Create the boolean mask for rows where 'words' contains 'apple' mask = df['words'].str.contains('apple') # Use the mask to filter the DataFrame and get the 'names' column, then convert to a list matching_names = df.loc[mask, 'names'].tolist() print(matching_names) # Output: ['a', 'b']
If you need exact matches (not partial contains, e.g., avoiding matches for 'applepie'), replace str.contains with a direct equality check:
mask = df['words'] == 'apple'
2. Advanced: Map multiple keywords to their corresponding names
For your second request—getting independent names lists for each keyword—you can use a dictionary comprehension to loop through your keywords and build a map of keyword-to-names.
Example with multiple keywords:
target_keywords = ['apple', 'pear'] # Build a dictionary where each key is a keyword, value is the list of matching names keyword_name_map = { keyword: df.loc[df['words'].str.contains(keyword), 'names'].tolist() for keyword in target_keywords } print(keyword_name_map) # Output: {'apple': ['a', 'b'], 'pear': ['c']}
Again, swap str.contains with == if you need exact matches. You can also add case=False to str.contains if you want case-insensitive matching (e.g., matching 'Apple' or 'APPLE' too):
mask = df['words'].str.contains('apple', case=False)
内容的提问来源于stack exchange,提问作者Gazoo

