基于块行的字母排序算法需求:按Machine names排序并保留模块行组
Got it, let's work through how to sort your machine data alphabetically by Machine names while keeping each machine's associated module rows intact as a single block. This is a common scenario when dealing with grouped data, so here are practical solutions using different tools:
Key Principles First
The core idea is to treat each machine and its linked module rows as a single "block" before sorting. We first identify these blocks, associate them with their machine names, sort the blocks by name, then output the sorted result.
Step-by-Step Breakdown
- Identify block boundaries: Each block starts with a
Machine names:line, followed by a#modules:line that tells you how many module rows come next. - Group each complete block with its corresponding machine name.
- Sort the grouped blocks alphabetically by machine name.
- Output the sorted blocks in order.
Python Solution
This script is flexible and easy to adapt if your data has slight variations:
def sort_machine_blocks(input_data): blocks = [] current_block = [] module_count = 0 lines = input_data.strip().split('\n') for line in lines: if line.startswith('Machine names:'): # Save the previous block if we're starting a new one if current_block: blocks.append(current_block) current_block = [line] elif line.startswith('#modules:'): module_count = int(line.split(':')[1].strip()) current_block.append(line) else: current_block.append(line) # Check if we've reached the end of the current block if len(current_block) - 2 == module_count: blocks.append(current_block) current_block = [] module_count = 0 # Catch the last block if it doesn't end exactly at module count if current_block: blocks.append(current_block) # Sort blocks by the machine name (strip whitespace for clean sorting) blocks.sort(key=lambda block: block[0].split(':')[1].strip()) # Combine blocks back into a single string for output return '\n\n'.join(['\n'.join(block) for block in blocks]) # Example usage with sample data raw_data = """Machine names: Bravo #modules: 3 Module1: baz Module2: qux Module3: quux Machine names: Alpha #modules: 2 Module1: foo Module2: bar Machine names: Charlie #modules: 1 Module1: corge""" print(sort_machine_blocks(raw_data))
When you run this, the output will be sorted alphabetically by machine name, with each complete block preserved:
Machine names: Alpha #modules: 2 Module1: foo Module2: bar Machine names: Bravo #modules: 3 Module1: baz Module2: qux Module3: quux Machine names: Charlie #modules: 1 Module1: corge
Command-Line (awk) Solution
If you prefer a terminal-based approach, awk is perfect for this kind of text processing:
BEGIN { # Set record separator to "Machine names:" to split blocks RS = "Machine names:" # Split each record into lines FS = "\n" } # Skip the first empty record created by RS NR > 1 { # Extract the machine name (first part of the record) machine_name = $1 # Rebuild the full block with the "Machine names:" prefix full_block = "Machine names:" $0 # Store the block in an array indexed by machine name blocks[machine_name] = full_block } END { # Sort the array keys alphabetically PROCINFO["sorted_in"] = "@ind_str_asc" # Print each sorted block with a blank line separator for (name in blocks) { print blocks[name] "\n" } }
How to Use
- Save the script as
sort_machines.awk - Run it against your input file:
awk -f sort_machines.awk your_input_file.txt
Notes to Avoid Issues
- Ensure each
Machine names:line is immediately followed by a#modules:line with the correct count of module rows. - If your data uses blank lines to separate blocks, adjust the
RS(record separator) in the awk script to"\n\n"instead.
内容的提问来源于stack exchange,提问作者M.Vnh

