如何在Pandas表格中根据指定Series拆分Tag列并展开行?
Got it, let's work through this problem step by step. We need to split the Tag column using a given vocabulary list, keep any non-vocabulary character chunks, and then expand each chunk into its own row with the original ID. Here's how to make it happen:
Step 1: Set up the sample data
First, let's recreate your input data to make this solution reproducible:
import pandas as pd import re # Original DataFrame df = pd.DataFrame({ 'ID': [1, 2, 3, 4], 'Tag': ['up down', 'right-left', 'summerfivehigh', 'up a right b left'] }) # Predefined vocabulary Series vocab = pd.Series(['up', 'down', 'right', 'left', 'summer'])
Step 2: Define the splitting logic
The core idea is to mark all vocabulary terms in the Tag string with a temporary delimiter, split the string on that delimiter, clean up unwanted whitespace/separators, and retain valid chunks.
We'll use regex to target vocabulary terms, wrap them in a delimiter, then process the split results:
# Create a regex pattern that matches any term in the vocabulary (escaped to handle special characters) vocab_pattern = re.compile('|'.join(map(re.escape, vocab))) def split_tag_string(tag): # Wrap vocabulary terms with a temporary '|' delimiter to isolate them tagged_str = vocab_pattern.sub(r'|\g<0>|', tag) # Split on the delimiter, strip extra spaces/hyphens, and filter out empty strings split_parts = [part.strip('- ') for part in tagged_str.split('|')] # Remove any leftover empty entries from splitting cleaned_parts = [part for part in split_parts if part] return cleaned_parts
Step 3: Apply the split and expand rows
Use assign() to apply our splitting function to the Tag column, then explode() to turn each list of chunks into individual rows:
# Apply the split function and expand the results into rows result_df = df.assign(Tag=df['Tag'].apply(split_tag_string)).explode('Tag').reset_index(drop=True) print(result_df)
Step 4: Check the output
Running this code will produce exactly the desired result:
ID Tag 0 1 up 1 1 down 2 2 right 3 2 left 4 3 summer 5 3 fivehigh 6 4 up 7 4 a 8 4 right 9 4 b 10 4 left
This approach handles all your edge cases: space-separated tags, hyphen-separated tags, concatenated terms like summerfivehigh, and non-vocabulary terms like a and b.
内容的提问来源于stack exchange,提问作者NBC

