编写Pig语句按性别统计评分记录数及平均评分并排查代码报错
Fixing Your Pig Script & HDFS Read Errors
Hey there, let's work through your issues one by one—first resolving the HDFS read failures, then tweaking your script to meet your full requirement (counting rating records AND calculating average ratings by gender).
1. Resolving HDFS Read Failures & IPC Port Error
The error messages you're seeing point to two main issues:
- Incorrect file paths in your Pig script: Your script uses relative paths like
.mike/users.txt, but the error shows Hadoop is looking for files athdfs://localhost:9000/mike/users.txtandhdfs://localhost:9000/mike/movies/data.txt. Relative paths in Pig resolve to your HDFS home directory (e.g.,/user/your_username/.mike/...), which doesn't match where your files are stored. - Wrong port for HDFS web access: The message about making an HTTP request to a Hadoop IPC port means you're trying to access
localhost:9000(HDFS's internal communication port) via a browser. The correct HDFS web UI port is50070—use that if you want to browse HDFS files visually.
Fix Steps:
- Use absolute HDFS paths in your script: Update the
LOADstatements to point directly to the correct file locations (match what the error message shows):- For users.txt:
/mike/users.txt - For data.txt:
/mike/movies/data.txt
- For users.txt:
- Verify files exist in HDFS: Run these commands in your terminal to confirm:
hdfs dfs -ls /mike/ hdfs dfs -ls /mike/movies/ - Check Hadoop cluster status: Ensure all necessary daemons are running with
jps—you should seeNameNode,DataNode,ResourceManager, andNodeManagerlisted.
2. Tweaking Your Pig Script to Meet the Full Requirement
Your original script only calculates average ratings, but you also need the count of rating records per gender. Here's the corrected, complete script with absolute paths and both metrics:
-- Load user data with absolute HDFS path A = LOAD '/mike/users.txt' USING PigStorage() AS (userID: int, age: int, gender: chararray, occupation: chararray, zip: int); -- Load rating data with correct absolute path B = LOAD '/mike/movies/data.txt' USING PigStorage() AS (userID: int, itemid: chararray, rating: int, timestamp: int); -- Join user and rating datasets on userID joinAB = JOIN A BY userID, B BY userID; -- Group the joined data by gender C = GROUP joinAB BY gender; -- Generate final output: gender, number of rating records, average rating D = FOREACH C GENERATE group AS gender, COUNT(joinAB) AS total_rating_records, AVG(joinAB.rating) AS average_rating; -- Dump the result to console DUMP D;
Key Changes:
- Added
COUNT(joinAB)to get the total number of rating records per gender - Used clear aliases (
total_rating_records,average_rating) for readability - Switched to absolute paths to avoid HDFS location confusion
Quick Testing Tip
If you want to verify each step works, run small test commands first:
- Load just the user data and dump it to confirm read access:
A = LOAD '/mike/users.txt' USING PigStorage() AS (userID: int, age: int, gender: chararray, occupation: chararray, zip: int); DUMP A; - Do the same for the rating data before moving on to joins and grouping.
内容的提问来源于stack exchange,提问作者hyena
相关产品推荐
相关产品推荐

