STORE_PRIMARYINPUTSUNAVAILABLEBEYONDLIMIT报错及ADLA文件数量限制咨询
Hey there, let's unpack what you're encountering with that ADLA error and your question about input file limits.
That STORE_PRIMARYINPUTSUNAVAILABLEBEYONDLIMIT error is directly tied to ADLA's handling thresholds for job input files. When processing a large number of small files in a single job, ADLA can hit limits related to underlying storage metadata read pressure, job scheduling resource allocation, and input processing overhead. The threshold isn't a hard, fixed number—it can vary based on file sizes, current cluster load, and storage performance. Your 242 files just happened to cross that threshold for your specific environment, while splitting the time window reduced the file count enough to stay under it.
Your solution of merging files is actually one of the most effective and recommended fixes here. Reducing the number of small files not only avoids this error but also boosts overall job execution efficiency by cutting down on metadata overhead and resource contention.
If you run into this again in the future, you can also try these additional tweaks:
- Adjust job parallelism settings: Increasing the number of parallel tasks can help ADLA allocate resources more efficiently to handle input files
- Verify storage account performance: Ensure your storage account isn't facing bandwidth or metadata access bottlenecks, which can exacerbate input processing limits
内容的提问来源于stack exchange,提问作者Tysklind-MSFT

