基于列值对Pandas DataFrame的Description列动态截取子串
动态截取DataFrame字符串列的实现方案
问题背景
给定如下Pandas DataFrame:
import pandas as pd data = {'Description': ['with lemon', 'lemon', 'and orange', 'orange'], 'Start': ['6', '1', '5', '1'], 'Length': ['5', '5', '6', '6']} df = pd.DataFrame(data) print(df)
需要根据Start和Length列的数值,对Description列做动态子串截取,最终得到包含Res结果列的DataFrame:
data = {'Description': ['with lemon', 'lemon', 'and orange', 'orange'], 'Start': ['6', '1', '5', '1'], 'Length': ['5', '5', '6', '6'], 'Res': ['lemon', 'lemon', 'orange', 'orange']} df = pd.DataFrame(data) print(df)
目前仅能实现固定位置的字符串截取(如df['Res'] = df['Description'].str[1:2]),需要更简洁的动态实现方式。
可行方案
方案一:使用apply逐行处理
先将Start和Length列转换为整数类型(原始数据为字符串格式),再通过apply遍历每一行,根据对应位置截取子串:
# 转换列类型为整数 df['Start'] = df['Start'].astype(int) df['Length'] = df['Length'].astype(int) # 生成结果列 df['Res'] = df.apply(lambda row: row['Description'][row['Start']-1 : row['Start']-1 + row['Length']], axis=1)
注意:Python字符串索引从0开始,示例中的Start是1-based计数,因此需要减1转换为0-based索引
方案二:使用str.slice批量处理
Pandas的str.slice方法支持直接传入列作为参数,实现高效的批量动态截取,代码更简洁且性能更优:
# 转换列类型为整数 df['Start'] = df['Start'].astype(int) df['Length'] = df['Length'].astype(int) # 动态截取子串 df['Res'] = df['Description'].str.slice(start=df['Start']-1, stop=df['Start']-1 + df['Length'])
该方法避免了逐行遍历,适合处理大规模数据集。
内容的提问来源于stack exchange,提问作者highbury
相关产品推荐
相关产品推荐

