如何修复Hive on MR执行SQL时的src filesystem资源变更问题
Hey there, let's break down the frustrating issue you're hitting with Hive on MapReduce. That error about the libjars resource changing mid-download is tricky, especially since you already synced NTP clocks. Let's walk through targeted fixes and checks you can try:
1. Tweak S3A Client Consistency Settings
Since you're using S3 as the storage backend, its eventual consistency model might be causing this metadata mismatch. Try adjusting these Hadoop configuration parameters:
- Increase
fs.s3a.consistent.retry.intervalto a higher value (like 5000ms) to give the storage time to sync metadata - Enable
fs.s3a.consistentif your S3 bucket uses a strongly consistent storage class (like S3 Standard) - Bump up
fs.s3a.download.retry.maxto let the client retry the download a few more times before failing
2. Lock Down the Staging Directory
Double-check that no other scripts, processes, or users are modifying the staging libjars directory (s3a://tpcds/tmp/hadoop-yarn/staging/root/.staging/job_1570514228864_0001/libjars) after you submit the Hive job. YARN captures resource metadata at submission time—any changes to that directory mid-job will trigger this exact error. Keep hands off that staging path until the job completes.
3. Check for Hadoop Version Bugs
Some older Hadoop versions have known bugs in the S3A client that mess up file modification time detection. Take a look at the release notes for your Hadoop version to see if there's a patch or upgrade that addresses this specific issue. Upgrading to a more recent stable release might resolve the metadata mismatch.
4. Manually Specify Libjars Instead of Auto-Staging
Instead of relying on YARN to auto-stage your libjars, try explicitly passing them with the --libjars flag when launching your Hive job. This skips the auto-staging process that might be introducing inconsistencies. Here's an example command:
hive --libjars /local/path/to/your/jars/*.jar -e "INSERT INTO your_table SELECT * FROM source_table;"
5. Verify S3 Cache Layers (If Applicable)
If you have a caching layer (like a CDN or internal cache) in front of your S3 bucket, it might be serving stale metadata to the YARN client. Try disabling the cache temporarily or forcing a refresh to ensure YARN gets the latest, accurate metadata for the libjars.
Hopefully one of these steps gets your Hive jobs running smoothly. Let us know if you need more details on any of these checks!
内容的提问来源于stack exchange,提问作者Eugene

