如何将指定语句存入字典?处理文件与数据集编号的技术问询
Let’s break this down into two straightforward steps: first pulling out the X-Y number pairs from each line of your dataset, then storing those pairs (and their associated data) in a dictionary for whatever processing you need next.
Step 1: Extract the X-Y Number Pairs
Your dataset lines follow a consistent path pattern: values/test/[X]/blueprint-[Y].png,.... To grab the X-Y pair, we can parse the path string directly. Here’s a Python example that works with your sample data:
def extract_number_pair(line): # Split the line to isolate the image path (first element before the comma) image_path = line.strip().split(',')[0] # Split the path into segments using '/' as the delimiter path_segments = image_path.split('/') # Get X: the third segment (index 2, since we start counting from 0) x = path_segments[2] # Get Y: pull from the blueprint filename (e.g., "blueprint-0.png" → "0") blueprint_filename = path_segments[3] y = blueprint_filename.split('-')[1].split('.')[0] # Return the formatted pair return f"{x}-{y}" # Test with your sample lines sample_lines = [ "values/test/10/blueprint-0.png,2089.0,545.0,2100.0,546.0", "values/test/10/blueprint-0.png,2112.0,545.0,2136.0,554.0", "values/test/45/blueprint-1.png,112.0,45.0,36.0,654.0" ] for line in sample_lines: print(extract_number_pair(line)) # Output: 10-0, 10-0, 45-1
Step 2: Store Pairs in a Dictionary
How you structure the dictionary depends on your end goal. Here are two common use cases:
Use Case 1: Count Occurrences of Each Pair
If you want to track how many times each X-Y pair appears in your dataset:
pair_counts = {} for line in sample_lines: pair = extract_number_pair(line) # Increment count if the pair exists, else set to 1 pair_counts[pair] = pair_counts.get(pair, 0) + 1 print(pair_counts) # Output: {'10-0': 2, '45-1': 1}
Use Case 2: Store All Associated Data for Each Pair
If you need to keep the numerical values linked to each pair (for averaging, filtering, etc.):
pair_data = {} for line in sample_lines: parts = line.strip().split(',') pair = extract_number_pair(line) # Convert the numerical values to floats (adjust if you need integers) values = [float(val) for val in parts[1:]] # Add to the dictionary: append values to the pair's list if pair not in pair_data: pair_data[pair] = [] pair_data[pair].append(values) print(pair_data) # Output: # { # '10-0': [[2089.0, 545.0, 2100.0, 546.0], [2112.0, 545.0, 2136.0, 554.0]], # '45-1': [[112.0, 45.0, 36.0, 654.0]] # }
Scalability Tips for Large Datasets
Since you have thousands of lines, keep these in mind:
- Process the file line-by-line instead of loading the whole thing into memory:
with open('your_dataset.txt', 'r') as f: for line in f: pair = extract_number_pair(line) # Update your dictionary here - Add error handling (like
try-exceptblocks) to skip or log malformed lines that don’t follow the expected pattern.
内容的提问来源于stack exchange,提问作者Jess

