You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

加载大型JSON数组文件至Azure Data Lake Analytics (ADLA)遇阻求助

Fixing ADLA Stuck Loading Large JSON Array Files

Hey there, let's dig into this frustrating issue you're hitting with Azure Data Lake Analytics (ADLA) loading your massive 110GB JSON array file (150 million rows). Stuck at 0% progress with no updates to files like algebra.xml is a classic sign of a single-vertex bottleneck—here's how to break through it:

Why This Happens

ADLA's default JSON handling treats a full JSON array as a single input entity. That means the entire 110GB file gets assigned to one single vertex to process, which is way beyond what a single node can handle. No parallelism = no progress, hence the stuck job.

Practical Solutions to Fix It

1. Convert JSON Array to Line-Delimited JSON (LDJ)

The biggest win is reformatting your file so each JSON object sits on its own line, like this:

{"id": 1, "data": "foo"}
{"id": 2, "data": "bar"}

Instead of the array format:

[{"id":1,"data":"foo"},{"id":2,"data":"bar"}]
  • How to do this: Use Azure Data Factory (ADF) Mapping Data Flows to parse the array and write out line-delimited JSON, or use a U-SQL custom extractor to handle the conversion directly in ADLA (no need to move the file twice).

2. Use a Custom U-SQL Extractor for Parallel Processing

If you can't reformat the file upfront, use a U-SQL script that splits the JSON array into individual objects before parsing. Microsoft provides a sample JsonSplit function that does exactly this:

-- First, read the entire array as a single string
@rawInput =
    EXTRACT jsonContent string
    FROM "/your-datalake-path/large-file.json"
    USING Extractors.Text(delimiter: '\0'); -- Treats the whole file as one line

-- Split the array into individual JSON objects
@splitObjects =
    SELECT JsonFunctions.JsonTuple(obj) AS parsedData
    FROM @rawInput
    CROSS APPLY
        (SELECT * FROM
            (SELECT Microsoft.Analytics.Samples.FormattedJson.JsonSplit(jsonContent) AS obj) AS splitArray) AS individualObjs;

-- Now you can query @splitObjects normally
@result =
    SELECT parsedData["id"] AS Id, parsedData["data"] AS Data
    FROM @splitObjects;

OUTPUT @result TO "/output/parsed-data.csv"
USING Outputters.Csv();

This splits the array into chunks that ADLA can distribute across multiple vertices, finally unlocking parallel processing.

3. Tune ADLA Job Resources

Even with properly formatted data, make sure your job has enough resources to handle the load:

  • Increase Degree of Parallelism (DOP): When submitting the job, set a higher DOP (e.g., 32 or 64, based on your ADLA account quota) to allocate more vertices.
  • Adjust Memory/Core Allocation: If you're hitting memory limits, tweak MemoryPerNode and CorePerNode in the job configuration to give each vertex more resources.

4. Split the Large File into Smaller Chunks

If reformatting isn't an option, split the 110GB file into smaller, manageable chunks (e.g., 10GB each). ADLA can process multiple files in parallel, with each file assigned to multiple vertices.

  • Use ADF's Copy Activity with chunking enabled, or Azure CLI tools to split the file locally before uploading. Just ensure each chunk is a valid JSON structure (preferably line-delimited for best results).

5. Verify Storage Configuration

Double-check that your ADLA account and Data Lake Storage (Gen1/Gen2) are in the same region to avoid cross-region latency. For Gen2, confirm hierarchical namespace is enabled, and there are no permission locks or throughput limits on the storage account that might be throttling reads.


内容的提问来源于stack exchange,提问作者jatal

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 07:26:47