You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何通过正则或Pandas处理extractall结果,实现每行原文本聚合匹配?

Solution to Extract Aggregated Regex Group Matches per Row

First, let's break down what we need: retain every original row from your Series, with columns for each regex group showing either the captured content (aggregated if multiple matches exist in the same group for a row) or a missing value if no match was found for that group. Here are two straightforward approaches to achieve this:

Approach 1: Capture the First Match per Group

If you only care about the first occurrence of each group in a text, str.extract() is perfect—it returns a DataFrame directly mapped to your original Series rows, with columns for each named group:

import re
import pandas as pd

regex = r"(?P<adv>This)|(?P<noun>test)"
texts = ["This is a test", "Random stuff with no match", "This and This again", "test test test"]
series = pd.Series(texts)

# Extract first match for each group
result = series.str.extract(regex)
print(result)

Output:

adv  noun
0  This   NaN
1   NaN   NaN
2  This   NaN
3   NaN  test

Approach 2: Aggregate All Matches per Group

If you need to capture all occurrences of each group in a text (and combine them into a single value), we'll use str.extractall() to get all matches, aggregate by the original row index, then join back to the original Series to keep rows with no matches:

# Extract all matches (returns a multi-index DataFrame: original index + match number)
all_matches = series.str.extractall(regex)

# Aggregate matches per original row: join non-NaN values with commas (adjust logic as needed)
aggregated_matches = all_matches.groupby(level=0).agg(
    lambda x: ', '.join(x.dropna()) if not x.dropna().empty else pd.NA
)

# Join back to original series to include rows with no matches
final_result = series.to_frame(name="original_text").join(aggregated_matches)
print(final_result)

Output:

original_text         adv         noun
0        This is a test       This       <NA>
1  Random stuff with no match       <NA>       <NA>
2      This and This again  This, This       <NA>
3          test test test       <NA>  test, test, test

Key Notes:

  • Your existing regex works fine here—since groups are mutually exclusive, the alternation | correctly captures matches into the appropriate group. No regex modifications are needed.
  • For aggregation, tweak the lambda function to fit your needs: use x.dropna().iloc[0] to keep only the first match, len(x.dropna()) to count matches, or any other custom logic.
  • Using pd.NA ensures consistent missing value handling for string columns.

内容的提问来源于stack exchange,提问作者arnaud

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.07 17:32:49