如何拆分游戏流派拼接字符串并提取主流派?
拆分流派字符串并提取主流派
问题分析
你用split(',')无效是因为目标字符串里没有逗号分隔符。流派的核心规律是每个流派均以大写字母开头,部分流派包含连字符(如Action-Adventure),借助正则表达式就能精准处理。
拆分字符串的实现
用Python的re模块即可完成,提供两种实用方法:
方法1:直接提取所有流派(推荐)
import re text = "ActionAction-AdventureShooterStealth" genres = re.findall(r'[A-Z][^A-Z]*', text) print(genres) # 输出: ['Action', 'Action-Adventure', 'Shooter', 'Stealth']
正则说明:
[A-Z]:匹配每个流派的起始大写字母[^A-Z]*:匹配后续所有非大写字母的字符(含连字符、小写字母),直到遇到下一个大写字母(下一流派的开头)停止
方法2:拆分后处理
import re text = "ActionAction-AdventureShooterStealth" genres = re.split(r'(?=[A-Z])', text)[1:] # 移除拆分后开头的空字符串 print(genres) # 输出: ['Action', 'Action-Adventure', 'Shooter', 'Stealth']
正则说明:
(?=[A-Z]):正向预查大写字母的位置,在该位置拆分且不消耗字符
提取主流派
主流派是拆分后列表的第一个元素,直接取genres[0]即可。如果是处理DataFrame的列,可用apply批量处理:
import pandas as pd import re # 示例数据 df = pd.DataFrame({ 'Game': ['A', 'B', 'C', 'D'], 'Genres': ['ActionComedyAdventure', 'AdventureComedy', 'NarrationShooting', 'ActionAction-AdventureShooterStealth'] }) # 定义提取函数 def get_major_genre(genre_str): genres = re.findall(r'[A-Z][^A-Z]*', genre_str) return genres[0] if genres else None # 添加新列 df['Major Genre'] = df['Genres'].apply(get_major_genre) print(df)
输出结果:
Game Genres Major Genre 0 A ActionComedyAdventure Action 1 B AdventureComedy Adventure 2 C NarrationShooting Narration 3 D ActionAction-AdventureShooterStealth Action
内容的提问来源于stack exchange,提问作者Girish Sawant
相关产品推荐
相关产品推荐

