You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Pandas表格中根据指定Series拆分Tag列并展开行?

Split Pandas DataFrame Tag Column by Predefined Vocabulary and Expand Rows

Got it, let's work through this problem step by step. We need to split the Tag column using a given vocabulary list, keep any non-vocabulary character chunks, and then expand each chunk into its own row with the original ID. Here's how to make it happen:

Step 1: Set up the sample data

First, let's recreate your input data to make this solution reproducible:

import pandas as pd
import re

# Original DataFrame
df = pd.DataFrame({
    'ID': [1, 2, 3, 4],
    'Tag': ['up down', 'right-left', 'summerfivehigh', 'up a right b left']
})

# Predefined vocabulary Series
vocab = pd.Series(['up', 'down', 'right', 'left', 'summer'])

Step 2: Define the splitting logic

The core idea is to mark all vocabulary terms in the Tag string with a temporary delimiter, split the string on that delimiter, clean up unwanted whitespace/separators, and retain valid chunks.

We'll use regex to target vocabulary terms, wrap them in a delimiter, then process the split results:

# Create a regex pattern that matches any term in the vocabulary (escaped to handle special characters)
vocab_pattern = re.compile('|'.join(map(re.escape, vocab)))

def split_tag_string(tag):
    # Wrap vocabulary terms with a temporary '|' delimiter to isolate them
    tagged_str = vocab_pattern.sub(r'|\g<0>|', tag)
    # Split on the delimiter, strip extra spaces/hyphens, and filter out empty strings
    split_parts = [part.strip('- ') for part in tagged_str.split('|')]
    # Remove any leftover empty entries from splitting
    cleaned_parts = [part for part in split_parts if part]
    return cleaned_parts

Step 3: Apply the split and expand rows

Use assign() to apply our splitting function to the Tag column, then explode() to turn each list of chunks into individual rows:

# Apply the split function and expand the results into rows
result_df = df.assign(Tag=df['Tag'].apply(split_tag_string)).explode('Tag').reset_index(drop=True)

print(result_df)

Step 4: Check the output

Running this code will produce exactly the desired result:

ID       Tag
0   1        up
1   1      down
2   2     right
3   2      left
4   3    summer
5   3  fivehigh
6   4        up
7   4         a
8   4     right
9   4         b
10  4      left

This approach handles all your edge cases: space-separated tags, hyphen-separated tags, concatenated terms like summerfivehigh, and non-vocabulary terms like a and b.

内容的提问来源于stack exchange,提问作者NBC

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 06:54:23