You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

编写Pig语句按性别统计评分记录数及平均评分并排查代码报错

Fixing Your Pig Script & HDFS Read Errors

Hey there, let's work through your issues one by one—first resolving the HDFS read failures, then tweaking your script to meet your full requirement (counting rating records AND calculating average ratings by gender).

1. Resolving HDFS Read Failures & IPC Port Error

The error messages you're seeing point to two main issues:

  • Incorrect file paths in your Pig script: Your script uses relative paths like .mike/users.txt, but the error shows Hadoop is looking for files at hdfs://localhost:9000/mike/users.txt and hdfs://localhost:9000/mike/movies/data.txt. Relative paths in Pig resolve to your HDFS home directory (e.g., /user/your_username/.mike/...), which doesn't match where your files are stored.
  • Wrong port for HDFS web access: The message about making an HTTP request to a Hadoop IPC port means you're trying to access localhost:9000 (HDFS's internal communication port) via a browser. The correct HDFS web UI port is 50070—use that if you want to browse HDFS files visually.

Fix Steps:

  • Use absolute HDFS paths in your script: Update the LOAD statements to point directly to the correct file locations (match what the error message shows):
    • For users.txt: /mike/users.txt
    • For data.txt: /mike/movies/data.txt
  • Verify files exist in HDFS: Run these commands in your terminal to confirm:
    hdfs dfs -ls /mike/
    hdfs dfs -ls /mike/movies/
    
  • Check Hadoop cluster status: Ensure all necessary daemons are running with jps—you should see NameNode, DataNode, ResourceManager, and NodeManager listed.

2. Tweaking Your Pig Script to Meet the Full Requirement

Your original script only calculates average ratings, but you also need the count of rating records per gender. Here's the corrected, complete script with absolute paths and both metrics:

-- Load user data with absolute HDFS path
A = LOAD '/mike/users.txt' USING PigStorage() AS (userID: int, age: int, gender: chararray, occupation: chararray, zip: int);

-- Load rating data with correct absolute path
B = LOAD '/mike/movies/data.txt' USING PigStorage() AS (userID: int, itemid: chararray, rating: int, timestamp: int);

-- Join user and rating datasets on userID
joinAB = JOIN A BY userID, B BY userID;

-- Group the joined data by gender
C = GROUP joinAB BY gender;

-- Generate final output: gender, number of rating records, average rating
D = FOREACH C GENERATE 
    group AS gender,
    COUNT(joinAB) AS total_rating_records,
    AVG(joinAB.rating) AS average_rating;

-- Dump the result to console
DUMP D;

Key Changes:

  • Added COUNT(joinAB) to get the total number of rating records per gender
  • Used clear aliases (total_rating_records, average_rating) for readability
  • Switched to absolute paths to avoid HDFS location confusion

Quick Testing Tip

If you want to verify each step works, run small test commands first:

  1. Load just the user data and dump it to confirm read access:
    A = LOAD '/mike/users.txt' USING PigStorage() AS (userID: int, age: int, gender: chararray, occupation: chararray, zip: int);
    DUMP A;
    
  2. Do the same for the rating data before moving on to joins and grouping.

内容的提问来源于stack exchange,提问作者hyena

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.29 00:57:40