Pig处理Google N-grams数据集报错:未设置Schema Tuple无法生成代码
Hey there! That error pops up because Pig can't infer a clear schema for one of your relations—usually when a GENERATE step doesn't properly name all output fields. Let's break down the issue and fix your script.
The Root Cause
In your final roundto step, you're using the ROUND_TO function but haven't given its result a named alias. Pig requires explicit field names to build a valid schema for each relation, and without that, it can't generate the necessary execution code.
Corrected Code
Here's the adjusted script with the fix highlighted:
inp = LOAD 'link to file' AS (ngram:chararray, year:int, occurences:float, books:float); filter_input = FILTER inp BY (occurences >= 400) AND (books >= 8); groupinp = GROUP filter_input BY ngram; sum_occ = FOREACH groupinp GENERATE FLATTEN(group) as ngram, SUM(filter_input.occurences) / SUM(filter_input.books) AS ntry; // Added an alias for the ROUND_TO result and simplified field references roundto = FOREACH sum_occ GENERATE ngram, ROUND_TO(ntry, 2) AS rounded_ntry;
What Changed?
- Added
AS rounded_ntry: This gives the rounded value a clear, explicit field name, letting Pig define a complete schema for theroundtorelation. - Simplified field references: You don't need to prefix fields with the relation name (like
sum_occ.ngram) inside theFOREACH—since you're operating directly onsum_occ, Pig already knows to look for those fields in that relation.
Quick Debug Tip
If you run into schema issues again, use the DESCRIBE command to inspect the schema of any relation:
DESCRIBE sum_occ; // Will show (ngram:chararray, ntry:float) DESCRIBE roundto; // After the fix, will show (ngram:chararray, rounded_ntry:float)
This helps you quickly spot where a schema is missing or incomplete.
内容的提问来源于stack exchange,提问作者thegreatcoder

