HDInsight POC技术问询:已完成Blob数据上传,需验证MapReduce及Hive查询
Hey there! Great job knocking out the first phase of your HDInsight POC—getting test data up to Azure Blob via AzCopy and accessing it through Visual Studio is a fantastic foundation. Let’s break down how to validate HDInsight’s MapReduce capabilities and write Hive queries, step by step.
First, assuming you’ve already spun up an HDInsight cluster (if not, you can create one quickly via the Azure Portal linked to your storage account), here’s how to test MapReduce:
Connect to your cluster via SSH
This is the most direct way to run Hadoop commands. Use this command in your terminal, replacing<clustername>with your actual cluster name:ssh sshuser@<clustername>-ssh.azurehdinsight.netYou’ll use the SSH credentials you set up when creating the cluster.
Run a sample MapReduce job
HDInsight comes preloaded with example MapReduce jobs (like WordCount) to help you verify functionality. Run this command to execute the WordCount job on a sample text file included with the cluster:yarn jar /usr/hdp/current/hadoop-mapreduce-client/hadoop-mapreduce-examples.jar wordcount /example/data/gutenberg/davinci.txt /example/data/wordcount-outputOnce the job finishes, view the results with:
hdfs dfs -cat /example/data/wordcount-output/part-r-00000Use your own Blob-stored data
To test with the data you uploaded via AzCopy, replace the input path in the command with your Blob storage path (using thewasb://protocol). For example, if your data is in a container namedmycontainerunder storage accountmystorage, in a foldermydata:yarn jar /usr/hdp/current/hadoop-mapreduce-client/hadoop-mapreduce-examples.jar wordcount wasb://mycontainer@mystorage.blob.core.windows.net/mydata/input_files/ wasb://mycontainer@mystorage.blob.core.windows.net/mydata/wordcount_results/Just make sure the output path doesn’t already exist (Hadoop will throw an error if it does).
Hive lets you query structured/semi-structured data using SQL-like syntax—perfect for analyzing your Blob data. Here’s how to get started:
Access Hive
You have a few options here:- Hive CLI: Type
hivein your SSH session to enter the interactive Hive shell. - Visual Studio Data Lake Tools: Since you’re already using VS to access Blob storage, this is a great option. Connect to your HDInsight cluster via the Azure Explorer, then create a new Hive script file.
- Azure Portal Hive View: Navigate to your HDInsight cluster in the portal, then select "Hive View" under the "Cluster dashboards" section.
- Hive CLI: Type
Create an external Hive table linked to your Blob data
External tables keep your data in Blob storage (instead of moving it to HDFS), so deleting the table won’t erase your underlying data. For example, if your data is comma-separated (CSV), use this query (replace placeholders with your details):CREATE EXTERNAL TABLE IF NOT EXISTS my_test_data ( record_id STRING, customer_name STRING, purchase_amount DOUBLE ) ROW FORMAT DELIMITED FIELDS TERMINATED BY ',' STORED AS TEXTFILE LOCATION 'wasb://<containername>@<storageaccountname>.blob.core.windows.net/<your_data_folder>/';If your data has a header row, you can add
TBLPROPERTIES ("skip.header.line.count"="1")at the end to ignore it.Run basic Hive queries
Once the table is created, you can run standard SQL-like queries. For example:- View the first 10 records:
SELECT * FROM my_test_data LIMIT 10; - Calculate total purchases per customer:
SELECT customer_name, SUM(purchase_amount) as total_spent FROM my_test_data GROUP BY customer_name ORDER BY total_spent DESC;
- View the first 10 records:
Save query results to Blob storage
To export your query results back to Blob storage, use this syntax:INSERT OVERWRITE DIRECTORY 'wasb://<containername>@<storageaccountname>.blob.core.windows.net/<output_folder>' ROW FORMAT DELIMITED FIELDS TERMINATED BY ',' SELECT customer_name, SUM(purchase_amount) as total_spent FROM my_test_data GROUP BY customer_name;
Quick Tips
- If you use Visual Studio Data Lake Tools, you can monitor query progress in real time and easily view results in a grid format.
- Double-check that your HDInsight cluster has access to your Blob storage account—if you linked the storage during cluster creation, this is already set up. If not, you’ll need to add the storage account key to the cluster’s configuration.
内容的提问来源于stack exchange,提问作者RaviLobo

