Apache Pig实现非空结果时才写入文件的技术求助
Got it, I've run into this exact issue with Pig before—since Pig is a declarative language, it doesn't have native flow control to skip a STORE operation when a relation is empty. The workaround combines Pig to calculate the record count and a shell script to handle the conditional execution. Here's how to make it work:
Step 1: Modify Your Pig Script to Calculate Count and Temp Store
First, update your existing Pig script to:
- Keep your existing logic to process data into the
intermediate_deleterelation. - Calculate the number of records in
intermediate_deleteand store that count to a temporary location. - Store the
intermediate_deletedata to a temporary path (instead of your final output path).
Here's what that looks like in code:
-- Load your old/new roles data (keep your existing logic here) old_roles = LOAD 'path/to/old_roles' USING PigStorage(',') AS (id:chararray, role:chararray); new_roles = LOAD 'path/to/new_roles' USING PigStorage(',') AS (id:chararray, role:chararray); -- Your existing transformation/join logic to get intermediate_delete old_ids = FOREACH old_roles GENERATE id; new_ids = FOREACH new_roles GENERATE id; deleted_ids = JOIN old_ids BY id LEFT OUTER, new_ids BY id WHERE new_ids.id IS NULL; intermediate_delete = JOIN old_roles BY id, deleted_ids BY old_ids.id; -- Calculate record count for intermediate_delete count_delete = GROUP intermediate_delete ALL; count_result = FOREACH count_delete GENERATE COUNT(intermediate_delete) AS record_count; STORE count_result INTO '/tmp/delete_count' USING PigStorage('\t'); -- Store intermediate_delete to a temporary path STORE intermediate_delete INTO '/tmp/delete_temp' USING PigStorage(',');
Step 2: Shell Script to Handle Conditional Move
Create a shell script that reads the count from the temporary location, then decides whether to move the temp data to your final output path (or run a dedicated STORE script if you prefer that approach).
For HDFS environments, use this shell script:
#!/bin/bash # Run the Pig script to process data and generate temp files pig -x mapreduce -f your_pig_script.pig # Read the record count from the temp count file RECORD_COUNT=$(hadoop fs -cat /tmp/delete_count/part-r-00000 | awk '{print $1}') # Check if there are records to keep if [ "$RECORD_COUNT" -gt 0 ]; then echo "Found $RECORD_COUNT records to save. Moving to final output path..." hadoop fs -mv /tmp/delete_temp/* /path/to/final_delete_output/ else echo "No records in intermediate_delete. Skipping final store." fi # Clean up temporary files hadoop fs -rm -r /tmp/delete_count hadoop fs -rm -r /tmp/delete_temp
If you're running in local mode, replace the hadoop fs commands with standard shell commands (like cat, mv, rm -r).
Why Your Initial Count Attempt Didn't Work
Pig operates on declarative logic—you define what data transformations you want, but you can't add "if-then-else" logic to skip entire operations like STORE directly in the Pig script. The count gives you the information you need, but you have to use an external shell script to act on that count conditionally.
Bonus Tips
- Use Pig parameters to make temp paths configurable (e.g.,
-param TEMP_PATH=/tmp/delete_temp) so you don't hardcode paths. - Ensure your temp paths are unique (add a timestamp suffix if you're running multiple jobs) to avoid conflicts.
内容的提问来源于stack exchange,提问作者rp_prasanna

