如何在Azure Data Factory中为Blob容器内的TXT文件追加静态表头
Hey there! Let's walk through how to append that static header from a separate file to each TXT in your Azure Blob Storage container using Azure Data Factory (ADF). Here's a step-by-step breakdown that should get you sorted:
First, make sure you have an Azure Blob Storage Linked Service connected to your storage account. Then create three parameterized datasets (they'll make looping easier):
- Header File Dataset: Point this directly to your header TXT file (e.g.,
header.txt). Use theDelimitedTextformat, and set "First row as header" to No (since this file is just the header content itself). - Source TXT Dataset: Point this to your target container, and add a string parameter called
fileNameto dynamically reference each file. Again useDelimitedText, set "First row as header" to No (we'll add our own header). - Output TXT Dataset: Point this to your desired output location (can be the same container or a separate folder). Add a
fileNameparameter here too—you can even use a pattern like@concat('processed_', dataset().fileName)to rename output files if needed.
Let's put together the pipeline components to handle the header merge:
2.1 Grab the Header Content
Add a Lookup activity (name it FetchHeader) and configure it to use your Header File Dataset. Check the "First row only" option—since your header file is just one line of text, this will pull that line into a variable we can reuse later. You'll reference this content with an expression like @activity('FetchHeader').output.firstRow.Column1 (adjust Column1 if your dataset uses a different default column name).
2.2 List All Target TXT Files
Add a Get Metadata activity (name it ListSourceFiles) and point it to your target container (use the Source TXT Dataset without filling the fileName parameter). In the "Field list", check Child items—this will retrieve a list of all files in the container. We'll filter out the header file next.
2.3 Loop Through Each Source File
Add a For Each activity (name it ProcessEachFile) and set its "Items" property to this expression:
@filter(activity('ListSourceFiles').output.childItems, not(equals(item().name, 'header.txt')))
This filters out the header file so we only process the other TXTs. You can leave "Sequential" unchecked for parallel processing (great for lots of files) or check it if you want to avoid any potential conflicts.
2.4 Merge Header & File Content (The Key Step)
Inside the For Each loop, add a Data Flow activity (name it MergeHeaderAndContent). Here's how to configure the data flow:
Add Header as a Source:
- Create a Source using your Header File Dataset.
- Add a
Derived Columntransformation, create a new column calledcombined_data, and set its value to your header column (e.g.,Column1). - Add another Derived Column to add a sort key: create
row_orderwith a value of1(this ensures the header stays first).
Add Current Source File as a Second Source:
- Create another Source using your Source TXT Dataset, and pass the current loop's filename into the
fileNameparameter with@item().name. - If your source files have their own existing headers, go to the Source settings and set "Skip first n rows" to
1to remove them. - Add a
Derived Columntransformation: createcombined_databy concatenating all columns with your delimiter (e.g.,concat(Column1, ',', Column2)for comma-separated files; if it's a single column, just use the column name). - Add another Derived Column for the sort key: set
row_orderto2.
- Create another Source using your Source TXT Dataset, and pass the current loop's filename into the
Union & Sort the Data:
- Add a
Uniontransformation, connect both sources to it, and set the union mode to "ByName" to match thecombined_dataandrow_ordercolumns. - Add a
Sorttransformation, sort byrow_orderin ascending order—this ensures the header row is first.
- Add a
Sink the Merged Content:
- Add a Sink using your Output TXT Dataset, pass the current filename (or your custom output name) into the
fileNameparameter. - In the Sink settings, set "First row as header" to No (since we're already writing the header as the first row of data). Use your desired delimiter, and set "Cleanup sink" to Yes if you want to overwrite existing files.
- Add a Sink using your Output TXT Dataset, pass the current filename (or your custom output name) into the
Before running the full pipeline, test each component individually:
- Verify the Lookup activity correctly pulls your header.
- Check the Get Metadata activity lists all your target files (excluding the header).
- Do a test run of the Data Flow with a single file to ensure the header is appended correctly.
Once everything checks out, trigger the pipeline—it'll process all your TXT files and append the header to each one!
Hope this walkthrough helps you get the job done smoothly. If you hit any roadblocks with specific steps, feel free to ask for more details.
内容的提问来源于stack exchange,提问作者Pakard

