PIG中跨行业部门聚合Worker值失败(ERROR 1066)求助
Hey there! Let's figure out why you're hitting that ERROR 1066 and get your city-sector employee rankings working smoothly.
First, What's ERROR 1066 Telling You?
This error usually means Pig can't process the alias test you're trying to iterate over (like with DUMP or STORE). It almost always traces back to issues in earlier steps—either your data didn't load correctly, your schema is mismatched, or your aggregation logic has a syntax/structural flaw.
Let's Break Down the Likely Issues & Fixes
1. Mismatched Schema with Your Nested JSON
You're using Elephant Bird's JsonLoader('-nestedLoad') which handles nested JSON, but your schema (as (...)) is incomplete or doesn't match your actual JSON structure. This is a super common pitfall when loading nested data.
For example, if your sectorAnalysis.json looks like this:
{"city": "New York", "sector": "Finance", "employees": [{"id": 1}, {"id": 2}, ...]}
Your schema needs to explicitly define the nested bag/tuple structure:
bus_data = LOAD 'sectorAnalysis.json' USING com.twitter.elephantbird.pig.load.JsonLoader('-nestedLoad') AS (city:chararray, sector:chararray, employees:bag{t:tuple(id:int)});
If your schema uses the wrong field names, data types, or doesn't account for nested fields, Pig will fail to parse the data correctly—making all subsequent steps (like aggregation) useless.
First test: Run DUMP bus_data right after loading. If this throws an error, your schema or JSON file is the problem.
2. Incomplete or Incorrect Aggregation Logic
Your goal is to count employees per sector per city, but you only shared the LOAD step. Chances are your aggregation code (grouping, counting) has issues that break the test alias.
Here's a complete, working flow to get your desired output:
-- Step 1: Load data with correct schema (adjust to match your JSON!) bus_data = LOAD 'sectorAnalysis.json' USING com.twitter.elephantbird.pig.load.JsonLoader('-nestedLoad') AS (city:chararray, sector:chararray, employees:bag{t:tuple(id:int)}); -- Step 2: Group by city AND sector to aggregate per combination grouped_sectors = GROUP bus_data BY (city, sector); -- Step 3: Calculate employee count per group (this creates your `test` alias) test = FOREACH grouped_sectors GENERATE group.city AS city, group.sector AS sector, SIZE(bus_data.employees) AS employee_count; -- Use SIZE if you're counting a list of employees; if your JSON has a direct count field, use that instead -- Step 4: Optional (but what you want): Rank sectors by employee count within each city ranked_by_city = FOREACH (GROUP test BY city) GENERATE group AS city, ORDER test BY employee_count DESC AS ranked_sectors; -- View the final result DUMP ranked_by_city;
If you skipped grouping, misspelled a field name, or used the wrong aggregation function (like trying to count a non-bag field with SIZE), Pig can't generate a valid test alias.
3. Corrupted or Malformed JSON
Elephant Bird's JsonLoader is strict—even one malformed JSON record (missing commas, unclosed brackets, etc.) will cause the entire load to fail. Check your sectorAnalysis.json for syntax errors:
- Use a local JSON validator tool to scan the file
- Try loading a small subset of valid records first to rule out bad data
4. Typos or Alias Misnaming
Double-check that your final alias is indeed test and that all prior steps reference fields/aliases correctly. A tiny typo (like bus_dta instead of bus_data) will break the pipeline entirely.
Quick Troubleshooting Checklist
- Run
DUMP bus_dataimmediately after loading to confirm data is parsed correctly - Verify your schema matches every field (and nested structure) in your JSON
- Check for typos in
GROUP,FOREACH, or field names - Validate your JSON file for syntax errors
内容的提问来源于stack exchange,提问作者StormsEdge

