如何为数据集新增列:根据下方字符匹配填充对应值
实现方法
完全可以通过代码实现这个需求,以下是基于Python的具体方案:
核心逻辑
遍历数据集的每一行,判断当前行是否为数值行:
- 如果是数值行,检查下一行是否属于同一元素且为非数值字符:
- 满足条件则将该字符填入当前行的新列,并跳过下一行(避免重复处理)
- 不满足条件则新列留空
- 非数值行直接跳过(因为它们是对应上一行的标签,无需单独保留)
代码实现
def process_data(input_data): processed_rows = [] index = 0 total_rows = len(input_data) while index < total_rows: # 拆分当前行的元素和值 current_line = input_data[index].strip() if not current_line: index += 1 continue element, value = current_line.split() # 判断当前行是否为数值行 is_numeric = False try: float(value) is_numeric = True except ValueError: is_numeric = False if is_numeric: tag = "" # 检查下一行是否符合条件 if index + 1 < total_rows: next_line = input_data[index+1].strip() next_element, next_value = next_line.split() if next_element == element: # 判断下一行是否为非数值 try: float(next_value) except ValueError: tag = next_value index += 1 # 跳过已处理的标签行 processed_rows.append((element, value, tag)) index += 1 return processed_rows # 示例:处理内置数据 sample_data = [ "Nickel 10", "Nickel U", "Nickel 10", "Nickel U", "Nickel 10", "Nickel U", "Nickel 1.4", "Nickel J", "Nickel 10", "Nickel U", "Nickel 10", "Nickel U", "Nickel 10", "Nickel U", "Sodium 8.1", "Sodium 7.4", "Sodium 6.2", "Sodium 7.6", "Sodium 7.9", "Sodium 6.9", "Sodium 7.8", "Sodium 8.9", "Sodium 9", "Sodium 7.9", "Sodium 7", "Sodium R", "Sodium 8.4", "Sodium 7.7" ] # 执行处理并输出结果 result = process_data(sample_data) print("Element\tValue\tTag") for row in result: print(f"{row[0]}\t{row[1]}\t{row[2]}")
从文件读取并保存结果
如果你的数据保存在文本文件中,可以用以下代码替代示例数据部分:
# 读取输入文件 with open("input_data.txt", "r") as infile: input_lines = [line for line in infile] # 处理数据 processed_result = process_data(input_lines) # 保存结果到文件 with open("output_data.txt", "w") as outfile: outfile.write("Element\tValue\tTag\n") for line in processed_result: outfile.write(f"{line[0]}\t{line[1]}\t{line[2]}\n")
内容的提问来源于stack exchange,提问作者Rachel Gladstone
相关产品推荐
相关产品推荐

