如何生成DataFrame的season_new列:保留有效值并补全season空值
解决方法
核心逻辑
优先保留season列的非空值,仅当season为空时,从programme列提取季数。关键是先锁定season列的优先级,再用正则提取补充空值。
可执行代码
import pandas as pd # 构造示例数据集 dt = pd.DataFrame({ 'programme': ["grey's anatomy s1", "friends season 1", "grey's anatomy s2", "big bang theory s2", "big bang theory", "peaky blinders"], 'season': [None, 1, None, 2, 1, 1] }) # 1. 初始化新列为season列的值 dt['season_new'] = dt['season'] # 2. 从programme中提取季数的数字部分 # 正则匹配"s/season+数字"格式,仅捕获数字 extracted = dt['programme'].str.extract(r'(?:season\s?|s\s?)(\d+)', expand=False) # 3. 用提取结果填充season_new的空值,转为整数类型 dt['season_new'] = dt['season_new'].fillna(extracted).astype(int)
执行结果
| programme | season | season_new |
|---|---|---|
| grey's anatomy s1 | None | 1 |
| friends season 1 | 1 | 1 |
| grey's anatomy s2 | None | 2 |
| big bang theory s2 | 2 | 2 |
| big bang theory | 1 | 1 |
| peaky blinders | 1 | 1 |
细节说明
- 正则
(?:season\s?|s\s?)(\d+):用非捕获组(?:...)匹配两种前缀(season或s,允许带空格),(\d+)仅捕获数字部分,避免提取多余字符。 fillna仅作用于season_new的空值(即原season为空的行),严格遵循"优先用season列"的要求。astype(int)统一列类型,保证结果和原season列格式一致。
内容的提问来源于stack exchange,提问作者Amirul Affiq
相关产品推荐
相关产品推荐

