You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python Pandas:基于列数据长度的列拼接问题求助

解决Pandas按规则拼接生成column3的问题

完整代码实现

import pandas as pd
import numpy as np

# 示例数据
data = {'column1':['af28912368', 'Nan', '234671', 'asr61239'],
        'column2':[701, np.nan, 761, 312]}
df = pd.DataFrame(data)

# 1. 把字符串型'Nan'转为Pandas识别的缺失值
df['column1'] = df['column1'].replace('Nan', pd.NA)

# 2. 按规则处理column1
processed_col1 = np.where(
    df['column1'].isna(),
    pd.NA,
    np.where(
        df['column1'].str.len() > 8,
        df['column1'].str[-8:],
        df['column1'].str.ljust(8)  # 左对齐,右侧补空格至8位
    )
)

# 3. 拼接生成column3,严格遵循column1为NaN时结果也为NaN的规则
df['column3'] = np.where(
    df['column1'].isna(),
    pd.NA,
    processed_col1 + '_' + df['column2'].astype(str)
)

print(df)

运行结果

column1column2column3
af28912368701.08912368_701.0
NaN
234671761.0234671 _761.0
asr61239312.0asr61239_312.0

关键逻辑说明

  • 缺失值处理:原始数据中的'Nan'是字符串类型,必须转为Pandas的pd.NA才能用isna()准确判断缺失情况。
  • column1格式化:
    • 长度超过8位:直接截取最后8位
    • 长度≤8位:用ljust(8)左对齐并在右侧补空格,确保总长度为8位
    • 缺失值:保持缺失状态
  • 拼接规则:只有当column1非缺失时才执行拼接,否则column3直接设为缺失值,完全匹配需求。

如果需要去掉column2的小数后缀(比如把701.0转为701),可以把拼接时的df['column2'].astype(str)替换为:

df['column2'].apply(lambda x: str(int(x)) if pd.notna(x) else str(x))

内容的提问来源于stack exchange,提问作者Sathish Kumar

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.11 16:45:32