You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将指定语句存入字典?处理文件与数据集编号的技术问询

Solution: Extracting Number Pairs and Storing in a Dictionary

Let’s break this down into two straightforward steps: first pulling out the X-Y number pairs from each line of your dataset, then storing those pairs (and their associated data) in a dictionary for whatever processing you need next.

Step 1: Extract the X-Y Number Pairs

Your dataset lines follow a consistent path pattern: values/test/[X]/blueprint-[Y].png,.... To grab the X-Y pair, we can parse the path string directly. Here’s a Python example that works with your sample data:

def extract_number_pair(line):
    # Split the line to isolate the image path (first element before the comma)
    image_path = line.strip().split(',')[0]
    # Split the path into segments using '/' as the delimiter
    path_segments = image_path.split('/')
    # Get X: the third segment (index 2, since we start counting from 0)
    x = path_segments[2]
    # Get Y: pull from the blueprint filename (e.g., "blueprint-0.png" → "0")
    blueprint_filename = path_segments[3]
    y = blueprint_filename.split('-')[1].split('.')[0]
    # Return the formatted pair
    return f"{x}-{y}"

# Test with your sample lines
sample_lines = [
    "values/test/10/blueprint-0.png,2089.0,545.0,2100.0,546.0",
    "values/test/10/blueprint-0.png,2112.0,545.0,2136.0,554.0",
    "values/test/45/blueprint-1.png,112.0,45.0,36.0,654.0"
]

for line in sample_lines:
    print(extract_number_pair(line))  # Output: 10-0, 10-0, 45-1

Step 2: Store Pairs in a Dictionary

How you structure the dictionary depends on your end goal. Here are two common use cases:

Use Case 1: Count Occurrences of Each Pair

If you want to track how many times each X-Y pair appears in your dataset:

pair_counts = {}

for line in sample_lines:
    pair = extract_number_pair(line)
    # Increment count if the pair exists, else set to 1
    pair_counts[pair] = pair_counts.get(pair, 0) + 1

print(pair_counts)  # Output: {'10-0': 2, '45-1': 1}

Use Case 2: Store All Associated Data for Each Pair

If you need to keep the numerical values linked to each pair (for averaging, filtering, etc.):

pair_data = {}

for line in sample_lines:
    parts = line.strip().split(',')
    pair = extract_number_pair(line)
    # Convert the numerical values to floats (adjust if you need integers)
    values = [float(val) for val in parts[1:]]
    
    # Add to the dictionary: append values to the pair's list
    if pair not in pair_data:
        pair_data[pair] = []
    pair_data[pair].append(values)

print(pair_data)
# Output:
# {
#   '10-0': [[2089.0, 545.0, 2100.0, 546.0], [2112.0, 545.0, 2136.0, 554.0]],
#   '45-1': [[112.0, 45.0, 36.0, 654.0]]
# }

Scalability Tips for Large Datasets

Since you have thousands of lines, keep these in mind:

  • Process the file line-by-line instead of loading the whole thing into memory:
    with open('your_dataset.txt', 'r') as f:
        for line in f:
            pair = extract_number_pair(line)
            # Update your dictionary here
    
  • Add error handling (like try-except blocks) to skip or log malformed lines that don’t follow the expected pattern.

内容的提问来源于stack exchange,提问作者Jess

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 04:19:21