新手求助:Python3按分隔符拆分大YAML文件的代码问题
Hey there! Let's work through your problem step by step—you're off to a great start with the incremental approach, so first let's fix the basic splitting logic, then we'll build up to using those IDs as filenames.
First: What's Wrong with Your Current Code?
Let's break down the key issues causing your code to fail:
- Incorrect line check: When you split the file content with
split('\n'), each line loses its trailing newline character. Soiwill never equal'-\n'—you need to check for just'-'instead. - Undefined variable: You tried to write
opto the output file, but that variable doesn't exist—you meant to writeholding. - Unclosed final entry: Your loop won't write the last YAML entry, since there's no trailing
'-'to trigger the write logic. - Redundant
close()calls: When using Python'swithstatement, files are automatically closed when the block ends, so you don't needmyfile.close()oropfile.close().
Fixed Basic Splitting Code (Counter-Based Filenames)
Here's a corrected version that will split your YAML file into numbered .md files:
with open('test-file.yml', 'r') as myfile: start = 0 cntr = 1 holding = '' # Iterate over each line directly instead of splitting first for line in myfile: stripped_line = line.strip() # Check if this is a YAML entry separator if stripped_line == '-': if start == 1: # Write the accumulated content to a file with open(f'{cntr}.md', 'w') as opfile: opfile.write(holding.strip() + '\n') holding = '' cntr += 1 else: start = 1 else: if start == 1: holding += line # Don't forget to write the last entry! if holding.strip(): with open(f'{cntr}.md', 'w') as opfile: opfile.write(holding.strip() + '\n')
Next Step: Using IDs as Filenames
To use the ID value from each YAML entry as the filename, it's better to parse the YAML properly instead of manually parsing lines (this avoids issues with varying YAML formatting). Here's how to do it:
1. Install PyYAML
First, install the PyYAML library to handle YAML parsing:
pip install pyyaml
2. ID-Based Splitting Code
This code will read the entire YAML file, extract each entry's ID, and write the entry to a file named [ID].md:
import yaml with open('test-file.yml', 'r') as myfile: # Load all YAML entries into a list yaml_entries = yaml.safe_load(myfile) for entry in yaml_entries: # Make sure the entry has an ID (add error handling if needed) if 'ID' not in entry: print(f"Skipping entry without ID: {entry}") continue entry_id = entry['ID'] # Convert the entry back to YAML string to write to file entry_yaml = yaml.dump(entry, sort_keys=False) with open(f'{entry_id}.md', 'w') as opfile: opfile.write(entry_yaml)
Key Tips & Best Practices
- Validate YAML first: Before splitting, make sure your large YAML files are valid. Invalid YAML will cause
yaml.safe_load()to fail. - Handle edge cases: Add error handling for entries without IDs, duplicate IDs (which would overwrite files), or very large entries.
- Test incrementally: Start with a small test file (like your
test-file.yml) to verify the code works before scaling to large files. - Preserve formatting: Using
yaml.dump()withsort_keys=Falseensures the order of keys in your YAML entries stays the same as the original file.
内容的提问来源于stack exchange,提问作者workerbee

