如何在Pandas中拆分特定行的结尾数字并对齐数据列
处理Pandas DataFrame中Test列结尾数字的拆分与右移问题
问题背景
我有一个Pandas DataFrame,其中部分行的Test列内容以数字结尾,需要拆分这些行的结尾数字并将其右移至右侧对应列中。
示例数据
构造DataFrame代码
check_df = pd.DataFrame({ 'Test': ["Absolute Neutrophil Count","Absolute Lymphocyte Count 2.9","Absolute Neutrophil Count"], 'Result1': [6.56,2.8,5.5], 'Result2': [5.14,2.6,4.8], 'Result3': [4.69,"10~9/L",5.2], 'Unit': ["10~9/L","1.0-3.0","10~9/L"], 'Range': ["4.0-10.0",None,"4.0-10.0"] })
原始数据展示
Test Result1 Result2 Result3 Unit Range 0 Absolute Neutrophil Count 6.56 5.14 4.69 10~9/L 4.0-10.0 1 Absolute Lymphocyte Count 2.9 2.80 2.60 10~9/L 1.0-3.0 None 2 Absolute Neutrophil Count 5.50 4.80 5.2 10~9/L 4.0-10.0
已尝试操作
我已用正则表达式识别出Test列中以数字结尾的行:
check_df.iloc[:,0].str.contains(r'\d+$')
期望结果
Test Result1 Result2 Result3 Unit Range 0 Absolute Neutrophil Count 6.56 5.14 4.69 10~9/L 4.0-10.0 1 Absolute Lymphocyte Count 2.9 2.80 2.60 10~9/L 1.0-3.0 2 Absolute Neutrophil Count 5.50 4.80 5.2 10~9/L 4.0-10.0
困惑点
不确定如何基于索引或其他方法拆分这些行,确保所有行/列数据保持正确的对齐表格格式。
解决方案
步骤1:拆分Test列,提取文本与数字
使用正则表达式拆分Test列,分离出纯文本部分和结尾的数字:
# 提取Test列的文本主体和结尾数字 split_result = check_df['Test'].str.extract(r'^(.*?)\s+(\d+\.\d+)$') # 更新Test列为纯文本部分(仅针对有结尾数字的行) check_df.loc[split_result[0].notna(), 'Test'] = split_result[0] # 将提取的数字转为浮点数,方便后续填充 extracted_num = split_result[1].astype(float)
步骤2:列数据右移并填充数字
对需要处理的行,将Result1至Range的列依次右移一列,再把提取的数字填入Result1:
# 标记需要处理的行(即Test列曾有结尾数字的行) mask = split_result[0].notna() # 定义需要右移的列顺序 cols_to_shift = ['Result1', 'Result2', 'Result3', 'Unit', 'Range'] # 执行右移操作:将前一列的值赋值给后一列 check_df.loc[mask, cols_to_shift[1:]] = check_df.loc[mask, cols_to_shift[:-1]].values # 把提取的数字填充到Result1列 check_df.loc[mask, 'Result1'] = extracted_num[mask]
最终处理结果
Test Result1 Result2 Result3 Unit Range 0 Absolute Neutrophil Count 6.56 5.14 4.69 10~9/L 4.0-10.0 1 Absolute Lymphocyte Count 2.9 2.80 2.60 10~9/L 1.0-3.0 2 Absolute Neutrophil Count 5.50 4.80 5.2 10~9/L 4.0-10.0
内容的提问来源于stack exchange,提问作者ViSa
相关产品推荐
相关产品推荐

