关于使用Google Document AI提取医疗报告中关联型多组键值对(表格型检测结果)的技术咨询
Hey Neil, let's break down how to tackle this problem with Google Document AI.
First, you're right that the standard Form Parsing processor works great for single key-value pairs like patient name or age, but it falls short when you need to link multiple test names to their corresponding results—since it doesn't preserve that row-level association.
Here's what you can do to get that structured test result array you want:
Leverage the Table Extraction feature
Medical test result tables are perfect for this capability. Document AI will identify the table structure, extract each row, and preserve the link between cells in the same row. The API response will include table elements grouped by rows, so you can iterate through each row (skipping headers if needed) to map test names to their matching results directly.Adjust your Custom Processor setup (if you need finer control)
If you tried a custom processor before and lost the test name-value link, switch to row-level labeling. When configuring the processor, define labels for the "Test Name" column and "Test Value" column, then set it to recognize that these labels are paired within individual rows. This way, the processor will associate each test name with its corresponding value in the same row, instead of treating them as isolated repeated fields.Post-process the API response to shape your desired output
No matter which approach you use, a bit of post-processing will turn the raw response into your target JSON structure:- Pull single key-value pairs (like name and age) using standard form parsing logic.
- For table data, loop through each extracted row, pull the test name and value from the respective cells, and add each pair as an object to the
testsarray.
Here's a simplified pseudocode example of the post-processing step:
# Get processed document from Document AI API document = document_ai_process_response() patient_data = {} tests_list = [] # Extract single field values for entity in document.entities: if entity.type == "PATIENT_NAME": patient_data["name"] = entity.text.strip() elif entity.type == "PATIENT_AGE": patient_data["age"] = int(entity.text.strip()) # Extract table test results for table in document.tables: # Skip header row if needed for row in table.rows[1:]: test_name = row.cells[0].text.strip() test_value = row.cells[1].text.strip() tests_list.append({"testName": test_name, "testValue": test_value}) patient_data["tests"] = tests_list
This will output exactly the structure you're looking for:
{ "name": "Jon Doe", "age": 76, "tests": [ { "testName": "CRP", "testValue": "51" }, { "testName": "Creatinine", "testValue": "0.8" } ] }
The core fix here is leaning into Document AI's table handling—it's designed specifically to preserve the row-wise relationships in grid-like structures, which is exactly what you need for linked test results.
备注:内容来源于stack exchange,提问作者Neil

