如何在Azure Data Factory复制活动中将.gz格式文件转换为.json格式
Unpack .GZ Files to JSON During ADF Copy Activity
Absolutely, you can unpack those .gz files and save the inner JSON content directly as uncompressed .json files in Azure Data Factory—no extra tools required! This will fix your downstream ETL issues with compressed files. Here's a step-by-step breakdown:
1. Configure Your Source Dataset
- Create or edit your source dataset, selecting Binary as the dataset type (since .gz files are binary compressed objects).
- Point the dataset to the storage location where your .gz files are stored (e.g., Azure Blob Storage, AWS S3, etc.).
2. Set Up the Copy Activity Source
- In your copy activity's Source tab, select the binary dataset you just configured.
- Expand the Compression section, set Compression type to
GZip. ADF will automatically decompress the .gz files and read the underlying JSON content during the copy process.
3. Configure the Sink Dataset for JSON Output
- Create a sink dataset targeting your Azure Data Lake Storage, selecting JSON as the dataset type.
- Define your target container/folder path. To preserve the original filename (replacing .gz with .json), use dynamic content for the filename field. For example:
This will take a file named@concat(split(item().name, '.gz')[0], '.json')data.gzand output it asdata.json.
4. Adjust Sink Settings in the Copy Activity
- In the copy activity's Sink tab, select your JSON sink dataset.
- In the JSON format settings, choose the appropriate File pattern based on your JSON structure:
- Use
Array of objectsif your JSON is a single array of objects. - Use
Set of objects(also called "JSON Lines") if each line in the file is a separate JSON object.
- Use
- Ensure the Compression setting is set to
None—this guarantees the output files are uncompressed JSON.
Quick Notes
- If a single .gz file contains multiple JSON files (instead of one), you'll need to use an ADF Data Flow to split the content into separate files. But for the common case of one JSON per .gz, the above steps work perfectly.
- Always test with a single .gz file first to verify the output JSON content and filename are correct before scaling to full batches.
内容的提问来源于stack exchange,提问作者Samyak Jain
相关产品推荐
相关产品推荐

