使用Pandas/Numpy实现字符串匹配查询的方法求助
Hey there! I see you've been stuck on this Pandas/Numpy problem for a while—let's break it down and get you the solution you need.
First, let's start with a quick assumption about your DataFrame structure (since you didn't share it explicitly, this is a common setup matching your example):
import pandas as pd # Sample DataFrame matching your example df = pd.DataFrame({ 'INDEXED_NUMBER': [0, 1, 2, 3], 'WORDS': ['aaa', 'bbb', 'ccc', 'aaa'] })
1. Exact Match (Exact String Match)
If you need to find the INDEXED_NUMBER where WORDS is exactly 'aaa', the simplest way is using boolean indexing with Pandas:
# Get all matching INDEXED_NUMBER values matching_numbers = df[df['WORDS'] == 'aaa']['INDEXED_NUMBER'] # If you only need the first match (like the 0 in your example) first_match = matching_numbers.iloc[0] print(first_match) # Outputs: 0
Or use .loc for more explicit syntax:
first_match = df.loc[df['WORDS'] == 'aaa', 'INDEXED_NUMBER'].iloc[0]
If there are multiple matches and you want all of them, just use matching_numbers.values to get a NumPy array of results.
2. Partial Match (String Contains Substring)
If your use case requires matching rows where WORDS contains the specified string (e.g., searching 'aa' should return rows with 'aaa'), use Pandas' .str.contains() method:
# Get all INDEXED_NUMBER where WORDS includes 'aaa' matching_numbers = df[df['WORDS'].str.contains('aaa')]['INDEXED_NUMBER'] # Get the first match first_match = matching_numbers.iloc[0]
Add case=False if you want case-insensitive matching (e.g., match 'AAA' or 'AaA' too):
matching_numbers = df[df['WORDS'].str.contains('AAA', case=False)]['INDEXED_NUMBER']
3. NumPy-Based Approach
If you prefer using NumPy instead, here's how you can achieve the same result:
import numpy as np # Convert columns to NumPy arrays words_array = df['WORDS'].values indexed_numbers_array = df['INDEXED_NUMBER'].values # Find indices where WORDS equals 'aaa' match_positions = np.where(words_array == 'aaa')[0] # Get corresponding INDEXED_NUMBER values matching_numbers = indexed_numbers_array[match_positions] # First match first_match = matching_numbers[0]
Common Pitfalls to Avoid
- Handle missing values: If your
WORDScolumn has NaN values,.str.contains()will throw an error. Addna=Falseto skip NaNs:df[df['WORDS'].str.contains('aaa', na=False)] - Check for empty matches: If no rows match your search,
iloc[0]will fail. Add a quick check to avoid crashes:matches = df[df['WORDS'] == 'aaa'] if not matches.empty: first_match = matches['INDEXED_NUMBER'].iloc[0] else: print("No matching records found!")
All these methods work perfectly in your Anaconda Prompt/Jupyter Notebook environment—just copy the relevant code blocks and adjust them to your actual DataFrame.
内容的提问来源于stack exchange,提问作者Lieberta

