You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

HDInsight POC技术问询:已完成Blob数据上传,需验证MapReduce及Hive查询

Hey there! Great job knocking out the first phase of your HDInsight POC—getting test data up to Azure Blob via AzCopy and accessing it through Visual Studio is a fantastic foundation. Let’s break down how to validate HDInsight’s MapReduce capabilities and write Hive queries, step by step.

Validating HDInsight MapReduce Functionality

First, assuming you’ve already spun up an HDInsight cluster (if not, you can create one quickly via the Azure Portal linked to your storage account), here’s how to test MapReduce:

  • Connect to your cluster via SSH
    This is the most direct way to run Hadoop commands. Use this command in your terminal, replacing <clustername> with your actual cluster name:

    ssh sshuser@<clustername>-ssh.azurehdinsight.net
    

    You’ll use the SSH credentials you set up when creating the cluster.

  • Run a sample MapReduce job
    HDInsight comes preloaded with example MapReduce jobs (like WordCount) to help you verify functionality. Run this command to execute the WordCount job on a sample text file included with the cluster:

    yarn jar /usr/hdp/current/hadoop-mapreduce-client/hadoop-mapreduce-examples.jar wordcount /example/data/gutenberg/davinci.txt /example/data/wordcount-output
    

    Once the job finishes, view the results with:

    hdfs dfs -cat /example/data/wordcount-output/part-r-00000
    
  • Use your own Blob-stored data
    To test with the data you uploaded via AzCopy, replace the input path in the command with your Blob storage path (using the wasb:// protocol). For example, if your data is in a container named mycontainer under storage account mystorage, in a folder mydata:

    yarn jar /usr/hdp/current/hadoop-mapreduce-client/hadoop-mapreduce-examples.jar wordcount wasb://mycontainer@mystorage.blob.core.windows.net/mydata/input_files/ wasb://mycontainer@mystorage.blob.core.windows.net/mydata/wordcount_results/
    

    Just make sure the output path doesn’t already exist (Hadoop will throw an error if it does).

Writing and Running Hive Queries

Hive lets you query structured/semi-structured data using SQL-like syntax—perfect for analyzing your Blob data. Here’s how to get started:

  • Access Hive
    You have a few options here:

    • Hive CLI: Type hive in your SSH session to enter the interactive Hive shell.
    • Visual Studio Data Lake Tools: Since you’re already using VS to access Blob storage, this is a great option. Connect to your HDInsight cluster via the Azure Explorer, then create a new Hive script file.
    • Azure Portal Hive View: Navigate to your HDInsight cluster in the portal, then select "Hive View" under the "Cluster dashboards" section.
  • Create an external Hive table linked to your Blob data
    External tables keep your data in Blob storage (instead of moving it to HDFS), so deleting the table won’t erase your underlying data. For example, if your data is comma-separated (CSV), use this query (replace placeholders with your details):

    CREATE EXTERNAL TABLE IF NOT EXISTS my_test_data (
        record_id STRING,
        customer_name STRING,
        purchase_amount DOUBLE
    )
    ROW FORMAT DELIMITED FIELDS TERMINATED BY ','
    STORED AS TEXTFILE
    LOCATION 'wasb://<containername>@<storageaccountname>.blob.core.windows.net/<your_data_folder>/';
    

    If your data has a header row, you can add TBLPROPERTIES ("skip.header.line.count"="1") at the end to ignore it.

  • Run basic Hive queries
    Once the table is created, you can run standard SQL-like queries. For example:

    • View the first 10 records:
      SELECT * FROM my_test_data LIMIT 10;
      
    • Calculate total purchases per customer:
      SELECT customer_name, SUM(purchase_amount) as total_spent
      FROM my_test_data
      GROUP BY customer_name
      ORDER BY total_spent DESC;
      
  • Save query results to Blob storage
    To export your query results back to Blob storage, use this syntax:

    INSERT OVERWRITE DIRECTORY 'wasb://<containername>@<storageaccountname>.blob.core.windows.net/<output_folder>'
    ROW FORMAT DELIMITED FIELDS TERMINATED BY ','
    SELECT customer_name, SUM(purchase_amount) as total_spent
    FROM my_test_data
    GROUP BY customer_name;
    

Quick Tips

  • If you use Visual Studio Data Lake Tools, you can monitor query progress in real time and easily view results in a grid format.
  • Double-check that your HDInsight cluster has access to your Blob storage account—if you linked the storage during cluster creation, this is already set up. If not, you’ll need to add the storage account key to the cluster’s configuration.

内容的提问来源于stack exchange,提问作者RaviLobo

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.26 09:47:16