You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何为数据集新增列:根据下方字符匹配填充对应值

实现方法

完全可以通过代码实现这个需求,以下是基于Python的具体方案:

核心逻辑

遍历数据集的每一行,判断当前行是否为数值行:

  • 如果是数值行,检查下一行是否属于同一元素且为非数值字符:
    • 满足条件则将该字符填入当前行的新列,并跳过下一行(避免重复处理)
    • 不满足条件则新列留空
  • 非数值行直接跳过(因为它们是对应上一行的标签,无需单独保留)

代码实现

def process_data(input_data):
    processed_rows = []
    index = 0
    total_rows = len(input_data)
    
    while index < total_rows:
        # 拆分当前行的元素和值
        current_line = input_data[index].strip()
        if not current_line:
            index += 1
            continue
        element, value = current_line.split()
        
        # 判断当前行是否为数值行
        is_numeric = False
        try:
            float(value)
            is_numeric = True
        except ValueError:
            is_numeric = False
        
        if is_numeric:
            tag = ""
            # 检查下一行是否符合条件
            if index + 1 < total_rows:
                next_line = input_data[index+1].strip()
                next_element, next_value = next_line.split()
                if next_element == element:
                    # 判断下一行是否为非数值
                    try:
                        float(next_value)
                    except ValueError:
                        tag = next_value
                        index += 1  # 跳过已处理的标签行
            processed_rows.append((element, value, tag))
        index += 1
    
    return processed_rows

# 示例:处理内置数据
sample_data = [
    "Nickel  10",
    "Nickel  U",
    "Nickel  10",
    "Nickel  U",
    "Nickel  10",
    "Nickel  U",
    "Nickel  1.4",
    "Nickel  J",
    "Nickel  10",
    "Nickel  U",
    "Nickel  10",
    "Nickel  U",
    "Nickel  10",
    "Nickel  U",
    "Sodium  8.1",
    "Sodium  7.4",
    "Sodium  6.2",
    "Sodium  7.6",
    "Sodium  7.9",
    "Sodium  6.9",
    "Sodium  7.8",
    "Sodium  8.9",
    "Sodium  9",
    "Sodium  7.9",
    "Sodium  7",
    "Sodium  R",
    "Sodium  8.4",
    "Sodium  7.7"
]

# 执行处理并输出结果
result = process_data(sample_data)
print("Element\tValue\tTag")
for row in result:
    print(f"{row[0]}\t{row[1]}\t{row[2]}")

从文件读取并保存结果

如果你的数据保存在文本文件中,可以用以下代码替代示例数据部分:

# 读取输入文件
with open("input_data.txt", "r") as infile:
    input_lines = [line for line in infile]

# 处理数据
processed_result = process_data(input_lines)

# 保存结果到文件
with open("output_data.txt", "w") as outfile:
    outfile.write("Element\tValue\tTag\n")
    for line in processed_result:
        outfile.write(f"{line[0]}\t{line[1]}\t{line[2]}\n")

内容的提问来源于stack exchange,提问作者Rachel Gladstone

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.25 13:48:15