如何通过for/df.iterrows在Pandas DataFrame中查找含‘rs’的子串并存储
Hey there! Let's work through this problem to get you the exact result you need. First, let's recap what you're asking for: find all values in your DataFrame that contain the substring "rs", map each matching value to its row index, and store everything in a dictionary (like {'rscat': 1} for your sample data). You mentioned using df.iterrows() because you found it efficient for large datasets, so we'll start with that approach, then also cover a more efficient vectorized method (since iterrows() can actually be slow for very big data—more on that later!).
Using df.iterrows() as requested
Here's how to implement this with iterrows(): we'll loop through each row, check every value in the row for the "rs" substring, and populate our dictionary with matches.
import pandas as pd # Your sample DataFrame data = {'First Column Name': ['AAA', 'BBB'], 'Second Column Name': ['CCC', 'rscat'], } df = pd.DataFrame(data, columns=['First Column Name', 'Second Column Name']) # Initialize empty dictionary for results result_dict = {} # Iterate over each row (returns index and row data) for idx, row in df.iterrows(): # Check every value in the current row for value in row.values: # Make sure we're dealing with a string to avoid errors if isinstance(value, str) and 'rs' in value: # Map the matching value to its row index result_dict[value] = idx # If you need to keep ALL indices for duplicate values (e.g., same value in multiple rows), # replace the line above with this block: # if value not in result_dict: # result_dict[value] = [] # result_dict[value].append(idx) print(result_dict) # Output: {'rscat': 1}
Notes for this approach:
- We add a check for string types to prevent errors if your DataFrame has non-string values (like numbers).
- If the same "rs"-containing value appears in multiple rows, the default code will overwrite the index with the last occurrence. Use the commented block if you want to store all corresponding indices in a list instead.
More efficient vectorized approach (better for large datasets)
While you mentioned iterrows() felt efficient, vectorized operations in Pandas are almost always faster for large datasets—they leverage optimized C-based operations under the hood instead of slow Python-level loops. Here's how to do the same task with vectorized methods:
# Create a boolean mask marking which elements contain "rs" # First convert all elements to strings to handle non-string data mask = df.astype(str).apply(lambda col: col.str.contains('rs')) # Extract matching values along with their row indices matches = df[mask].stack().reset_index() # Convert to the desired dictionary format result_dict_vectorized = matches.set_index(0)['level_0'].to_dict() print(result_dict_vectorized) # Output: {'rscat': 1}
Handling duplicate values with vectorized code:
If you need to keep all indices for repeated values, use groupby to collect indices into a list:
matches_grouped = matches.groupby(0)['level_0'].apply(list).to_dict() # Example output if "rscat" was in rows 1 and 3: {'rscat': [1, 3]}
This method will scale much better with huge datasets compared to iterrows()—definitely give it a test if you're working with millions of rows!
内容的提问来源于stack exchange,提问作者user14618562

