如何在Apache NiFi中提取CSV文件的最新记录并转为JSON
提取CSV最新记录并转为JSON的实现方案
Python 实现
这种方式适合大部分场景,尤其是处理大文件时,通过文件指针定位最后一行,避免加载整个文件占用过多内存:
import csv import json def get_latest_csv_record(csv_file_path): with open(csv_file_path, 'r', newline='') as f: # 获取CSV表头,作为JSON的键名 reader = csv.DictReader(f) headers = reader.fieldnames # 定位到文件末尾,向前查找最后一条数据行 f.seek(0, 2) position = f.tell() while position > 0: position -= 1 f.seek(position) if f.read(1) == '\n': break # 读取最后一行数据 last_line = f.readline().strip() if not last_line: return None # 将行数据与表头配对成字典 latest_record = dict(zip(headers, last_line.split(','))) return latest_record # 调用示例 csv_file = "your_data.csv" latest_data = get_latest_csv_record(csv_file) if latest_data: with open("latest_record.json", "w") as outfile: json.dump(latest_data, outfile, indent=4)
- 优势:高效处理大文件,自动匹配表头作为JSON键,兼容仅含表头无数据的边界情况。
Shell 脚本实现(Linux/macOS)
如果需要快速自动化处理,可结合jq工具实现:
固定表头场景
假设CSV表头为name,age,timestamp:
tail -n 1 your_data.csv | awk -F ',' '{print "{\"name\":\""$1"\",\"age\":\""$2"\",\"timestamp\":\""$3"\"}"}' > latest_record.json
动态表头场景
自动适配任意表头格式:
# 获取表头和最后一行数据 HEADERS=$(head -n 1 your_data.csv | tr ',' '\n') VALUES=$(tail -n 1 your_data.csv | tr ',' '\n') # 配对表头与值并转为标准JSON paste -d ':' <(echo "$HEADERS") <(echo "$VALUES") | jq -n 'reduce inputs as $i ({}; .[$i|split(":")[0]] = $i|split(":")[1])' > latest_record.json
- 注意:需先安装
jq工具(Debian/Ubuntu用apt install jq,macOS用brew install jq)。
内容的提问来源于stack exchange,提问作者BHADRAKA HERATH
相关产品推荐
相关产品推荐

