You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何通过Shell脚本获取HDFS中竖线分隔文本首行第2字段

Hey there! Let's work through this problem together since you're new to Unix and HDFS commands.

First, let's fix that "unknown command" error you ran into with -head:

  • The hadoop fs command doesn't have a built-in -head subcommand, which is why you're seeing that error. Instead, use hadoop fs -cat to read the file content, then pipe it to your system's native head command to grab just the first line.

Next, since your file uses vertical bars (|) as separators, your original awk command won't work as-is—it defaults to splitting on whitespace. You need to explicitly tell awk to use | as the field delimiter with the -F'|' flag.

Putting it all together, here's the one-liner to get the second field from the first line and store it in a variable:

# Store the second pipe-separated field from the first line in a variable
second_field=$(hadoop fs -cat "$new_file" | head -n 1 | awk -F'|' '{print $2}')

# Verify the result
echo "$second_field"

Let me break down each part for clarity:

  • hadoop fs -cat "$new_file": Reads the content of your HDFS file and sends it to standard output
  • head -n 1: Filters out everything except the complete first line
  • awk -F'|' '{print $2}': Splits the first line using | as the separator, then prints the second field

A quick note: If your Hadoop environment uses hdfs dfs instead of hadoop fs (they're mostly interchangeable), you can swap those commands. Some newer HDFS versions have a hdfs dfs -head command, but it outputs the first 1KB of the file—not guaranteed to be a full line—so cat | head -n 1 is more reliable for getting the complete first line.

内容的提问来源于stack exchange,提问作者pooja

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.29 08:12:32