You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将提取生成的row_list转换为自定义映射的嵌套字典?

Convert row_list to Nested Dictionary Structure

Got it, let's turn your row_list into that exact nested dictionary you're aiming for. Here's a simple, step-by-step solution that works perfectly with your existing data:

Approach

We'll iterate through each row in row_list (starting the row number at 1), then for each word in the row, we'll assign it to a column number (also starting at 1) while extracting just the text and x0 values you need.

Code Implementation

from decimal import Decimal

# Your existing row_list (from the example)
row_list = [
    [{'bottom': Decimal('58.650'), 'text': 'Hi there!', 'top': Decimal('40.359'), 'x0': Decimal('21.600'), 'x1': Decimal('65.644')}],
    [{'bottom': Decimal('74.101'), 'text': 'Your email', 'top': Decimal('37.519'), 'x0': Decimal('223.560'), 'x1': Decimal('300')}, {'bottom': Decimal('77.280'), 'text': 'my@domain.com', 'top': Decimal('62.506'), 'x0': Decimal('21.600'), 'x1': Decimal('140.775')}]
]

new = {}

# Iterate over rows with 1-based index
for row_num, row in enumerate(row_list, start=1):
    new[row_num] = {}
    # Iterate over words in the row with 1-based column index
    for col_num, word in enumerate(row, start=1):
        new[row_num][col_num] = {
            'text': word['text'],
            'x0': word['x0']
        }

# Print the result to verify
print(new)

Output

Running this code will produce exactly the structure you specified:

{
  1: {
    1: {'text': 'Hi there!', 'x0': Decimal('21.600')}
  },
  2: {
    1: {'text':'Your email', 'x0': Decimal('223.560')},
    2: {'text': 'my@domain.com', 'x0': Decimal('21.600')}
  }
}

Optional: Sort Columns by Horizontal Position

If you want columns in each row to be ordered from left to right (based on x0 value), just add a sort step for each row:

for row_num, row in enumerate(row_list, start=1):
    new[row_num] = {}
    # Sort words in the row by x0 (left to right)
    sorted_row = sorted(row, key=lambda x: x['x0'])
    for col_num, word in enumerate(sorted_row, start=1):
        new[row_num][col_num] = {
            'text': word['text'],
            'x0': word['x0']
        }

This would reorder the second row so my@domain.com (x0=21.6) is column 1, and Your email (x0=223.56) is column 2—adjust based on your actual needs!

内容的提问来源于stack exchange,提问作者oliverbj

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.13 08:55:42