如何通过Shell脚本获取HDFS中竖线分隔文本首行第2字段
Hey there! Let's work through this problem together since you're new to Unix and HDFS commands.
First, let's fix that "unknown command" error you ran into with -head:
- The
hadoop fscommand doesn't have a built-in-headsubcommand, which is why you're seeing that error. Instead, usehadoop fs -catto read the file content, then pipe it to your system's nativeheadcommand to grab just the first line.
Next, since your file uses vertical bars (|) as separators, your original awk command won't work as-is—it defaults to splitting on whitespace. You need to explicitly tell awk to use | as the field delimiter with the -F'|' flag.
Putting it all together, here's the one-liner to get the second field from the first line and store it in a variable:
# Store the second pipe-separated field from the first line in a variable second_field=$(hadoop fs -cat "$new_file" | head -n 1 | awk -F'|' '{print $2}') # Verify the result echo "$second_field"
Let me break down each part for clarity:
hadoop fs -cat "$new_file": Reads the content of your HDFS file and sends it to standard outputhead -n 1: Filters out everything except the complete first lineawk -F'|' '{print $2}': Splits the first line using|as the separator, then prints the second field
A quick note: If your Hadoop environment uses hdfs dfs instead of hadoop fs (they're mostly interchangeable), you can swap those commands. Some newer HDFS versions have a hdfs dfs -head command, but it outputs the first 1KB of the file—not guaranteed to be a full line—so cat | head -n 1 is more reliable for getting the complete first line.
内容的提问来源于stack exchange,提问作者pooja

