如何使用Pandas拆分非列表格式的DataFrame行并转换特定结构的数据?
拆分Pandas DataFrame中含分隔符的行
针对你的需求,我们可以分步骤处理这个DataFrame,把带/的行拆分成多行,同时保留对应关联的数据。下面是具体的实现过程:
首先,先把你的原始DataFrame加载进来(假设你已经导入了pandas):
import pandas as pd # 构造你的原始DataFrame data = { 'a': [19, 40, 41, 112, 154, 2991, 2992, 3887, 3893, 3908], 'b': ['560 80', '4.5 95', '1.76 95', '0.17/0.43 >95/>95', '7.2/1 >95/>95', '55 95', '33 95', '6.1 87.7', '3.9 70.3', '100 40'] } df = pd.DataFrame(data)
步骤1:拆分b列为两个独立字段
你的b列其实包含了两个数据项(比如数值和对应的阈值),所以先按空格把b列拆成两列,命名为value和threshold:
df[['value', 'threshold']] = df['b'].str.split(' ', expand=True) # 此时可以删掉原来的b列,简化数据 df = df.drop('b', axis=1)
步骤2:拆分带/的行并展开成多行
接下来,对value和threshold列,用str.split('/')把带分隔符的内容转换成列表,然后用explode方法将列表展开成多行,同时保持a列的对应值:
# 先把两个列都转成列表格式 df['value'] = df['value'].str.split('/') df['threshold'] = df['threshold'].str.split('/') # 同时展开两个列的列表 df = df.explode(['value', 'threshold'], ignore_index=True)
步骤3:处理>95的阈值转换
根据你的例子,需要把>95转换成95,可以用str.replace来处理:
df['threshold'] = df['threshold'].str.replace('>95', '95') # 如果需要把数值转成float类型,可以加上这一步 df['value'] = df['value'].astype(float) df['threshold'] = df['threshold'].astype(float)
最终结果
运行完上面的代码后,你得到的DataFrame就和你想要的结构一致了:
| a | value | threshold | |
|---|---|---|---|
| 0 | 19 | 560.0 | 80.0 |
| 1 | 40 | 4.5 | 95.0 |
| 2 | 41 | 1.76 | 95.0 |
| 3 | 112 | 0.17 | 95.0 |
| 4 | 112 | 0.43 | 95.0 |
| 5 | 154 | 7.2 | 95.0 |
| 6 | 154 | 1.0 | 95.0 |
| ... | ... | ... | ... |
额外说明
- 如果你的原始DataFrame中还有其他类似的分隔符(比如除了
/之外的),可以调整str.split的参数来适配。 explode方法在Pandas 0.25及以上版本支持同时展开多个列,如果你用的是旧版本,可以先展开一个列,再处理另一个列。
内容的提问来源于stack exchange,提问作者A.Razavi
相关产品推荐
相关产品推荐

