如何在Python 3+中遍历重排278行字符串并生成目标字典
Let’s break down how to solve your problem step by step—addressing the regex misalignment, fixing missing leading zeros, and building the sorted dictionary you need.
Step 1: Refine the Regex to Avoid Breaking App Names
The issue with your original regex (\s\d) is that it’s too broad and targets any space followed by a digit, which will accidentally chop parts of app names that include numbers or spaces. Instead, we need to isolate the numeric ID from the app name using a pattern that accounts for consistent structure in your input.
Common Scenario 1: ID comes first (with possible missing leading zeros)
If each line starts with the app ID (even if it’s missing leading zeros) followed by the app name, use this regex to split them safely:
import re # Matches leading digits, then captures everything after as the app name pattern = r'^(\d+)\s+(.*)$' # Example usage for a line like "1 My Cool App" line = "1 My Cool App" match = re.match(pattern, line) if match: app_id = match.group(1).zfill(4) # Adds leading zeros to make IDs uniform (e.g., "1" → "0001") app_name = match.group(2).strip()
Common Scenario 2: ID comes at the end of the line
If your lines are formatted as [App Name] [ID] (like "My Cool App 123"), use a non-greedy match to avoid capturing spaces within the app name:
pattern = r'^(.*?)\s+(\d+)$' line = "My Cool App 123" match = re.match(pattern, line) if match: app_name = match.group(1).strip() app_id = match.group(2).zfill(4)
Step 2: Process All Lines & Build the Sorted Dictionary
Once you can reliably split each line into ID and name, loop through your file, collect the pairs, and sort them as needed:
app_dict = {} sorted_app_list = [] with open('your_app_file.txt', 'r') as f: for line in f: line = line.strip() if not line: continue # Skip empty lines # Use the pattern that matches your input structure match = re.match(r'^(\d+)\s+(.*)$', line) if match: app_id = match.group(1).zfill(4) app_name = match.group(2).strip() # Add to a dictionary (ID as key, name as value) app_dict[app_id] = app_name # Or collect in a list for sorting sorted_app_list.append((app_id, app_name)) # Sort the list numerically by ID sorted_app_list.sort(key=lambda x: int(x[0])) # Convert to a sorted dictionary (Python 3.7+ preserves insertion order) sorted_app_dict = {id: name for id, name in sorted_app_list}
Step 3: Handle Edge Cases
- App names with numbers: The non-greedy
.*?in Scenario 2 ensures we only capture the final numeric sequence as the ID, even if the app name includes numbers (e.g., "App 2024" won’t be split incorrectly). - Inconsistent spacing: Using
\s+(matches one or more spaces) instead of a single space handles varying gaps between the ID and name. - Missing leading zeros:
zfill(n)standardizes ID length—adjustnto match your desired format (e.g.,zfill(3)for 3-digit IDs).
内容的提问来源于stack exchange,提问作者tryin2code

