Python提取DataFrame中File_Path列的第4级目录路径
路径提取解决方案:提取路径树第4层级并处理特殊情况
原始数据
import pandas as pd df = pd.DataFrame({ 'File_Path': [ '/data/application/AANX/aanx-dataeng-slas-sysyphus/scripts/s_shell/call_iws/call_PP_NEXT_RTBA_MAU_IND_INVE_D.sh', '/data/application/AANX/aanx-dataeng-slas-sysyphus/scripts/s_shell/call_iws/call_PP_NEXT_RTBA_MAU_IND_EMPF_D.sh', '/data/app_next_best_action/call_nba_as11.sh', '/data/application/AAIN/aain-srv-motor-extracao-next/iws/call_run_extract_default.sh cdlc_ing', 'sh /data/processos/current/aplicacao/AAVR/ACN10/scr/exec_fim_grupo.sh ACN10_ARQ_1' ] })
需求说明
提取File_Path列路径树结构的第4项,特殊情况处理:
- 路径层级不足4层时,直接取整个路径
- 路径开头带有
sh前缀的,先去除该前缀再处理
实现代码
def extract_parent_path(file_path): # 移除开头的"sh "前缀 cleaned_path = file_path.replace(r'^sh\s+', '', regex=True) # 分离路径与后续参数,仅保留路径部分 path_part = cleaned_path.split(' ', 1)[0] # 分割路径为层级列表,过滤空元素(处理绝对路径开头的/) path_levels = [level for level in path_part.split('/') if level] if len(path_levels) >= 4: # 取前4层,拼接为带首尾/的路径 return '/' + '/'.join(path_levels[:4]) + '/' else: # 层级不足4层,返回完整路径 return path_part # 应用函数生成新列 df['Parent_path'] = df['File_Path'].apply(extract_parent_path)
处理结果
File_Path Parent_path 0 /data/application/AANX/aanx-dataeng-slas-sysyphus/scripts/s_shel... /data/application/AANX/aanx-dataeng-slas-sysyphus/ 1 /data/application/AANX/aanx-dataeng-slas-sysyphus/scripts/s_shel... /data/application/AANX/aanx-dataeng-slas-sysyphus/ 2 /data/app_next_best_action/call_nba_as11.sh /data/app_next_best_action/call_nba_as11.sh 3 /data/application/AAIN/aain-srv-motor-extracao-next/iws/call_run... /data/application/AAIN/aain-srv-motor-extracao-next/ 4 sh /data/processos/current/aplicacao/AAVR/ACN10/scr/exec_fim_gr... /data/processos/current/aplicacao/
内容的提问来源于stack exchange,提问作者gfernandes
相关产品推荐
相关产品推荐

