You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何通过for/df.iterrows在Pandas DataFrame中查找含‘rs’的子串并存储

Hey there! Let's work through this problem to get you the exact result you need. First, let's recap what you're asking for: find all values in your DataFrame that contain the substring "rs", map each matching value to its row index, and store everything in a dictionary (like {'rscat': 1} for your sample data). You mentioned using df.iterrows() because you found it efficient for large datasets, so we'll start with that approach, then also cover a more efficient vectorized method (since iterrows() can actually be slow for very big data—more on that later!).

Using df.iterrows() as requested

Here's how to implement this with iterrows(): we'll loop through each row, check every value in the row for the "rs" substring, and populate our dictionary with matches.

import pandas as pd

# Your sample DataFrame
data = {'First Column Name': ['AAA', 'BBB'], 'Second Column Name': ['CCC', 'rscat'], }
df = pd.DataFrame(data, columns=['First Column Name', 'Second Column Name'])

# Initialize empty dictionary for results
result_dict = {}

# Iterate over each row (returns index and row data)
for idx, row in df.iterrows():
    # Check every value in the current row
    for value in row.values:
        # Make sure we're dealing with a string to avoid errors
        if isinstance(value, str) and 'rs' in value:
            # Map the matching value to its row index
            result_dict[value] = idx
            # If you need to keep ALL indices for duplicate values (e.g., same value in multiple rows),
            # replace the line above with this block:
            # if value not in result_dict:
            #     result_dict[value] = []
            # result_dict[value].append(idx)

print(result_dict)
# Output: {'rscat': 1}

Notes for this approach:

  • We add a check for string types to prevent errors if your DataFrame has non-string values (like numbers).
  • If the same "rs"-containing value appears in multiple rows, the default code will overwrite the index with the last occurrence. Use the commented block if you want to store all corresponding indices in a list instead.

More efficient vectorized approach (better for large datasets)

While you mentioned iterrows() felt efficient, vectorized operations in Pandas are almost always faster for large datasets—they leverage optimized C-based operations under the hood instead of slow Python-level loops. Here's how to do the same task with vectorized methods:

# Create a boolean mask marking which elements contain "rs"
# First convert all elements to strings to handle non-string data
mask = df.astype(str).apply(lambda col: col.str.contains('rs'))

# Extract matching values along with their row indices
matches = df[mask].stack().reset_index()

# Convert to the desired dictionary format
result_dict_vectorized = matches.set_index(0)['level_0'].to_dict()

print(result_dict_vectorized)
# Output: {'rscat': 1}

Handling duplicate values with vectorized code:

If you need to keep all indices for repeated values, use groupby to collect indices into a list:

matches_grouped = matches.groupby(0)['level_0'].apply(list).to_dict()
# Example output if "rscat" was in rows 1 and 3: {'rscat': [1, 3]}

This method will scale much better with huge datasets compared to iterrows()—definitely give it a test if you're working with millions of rows!

内容的提问来源于stack exchange,提问作者user14618562

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.11 07:36:19