基于Python 3实现todo.txt:从字符串提取待办任务信息至字典
Python 3 Todo.txt Parser Function
Got it, let's build a robust function to parse todo.txt entries into a structured dictionary. I’ve worked with this format extensively, so I’ll break down each component step by step to handle all edge cases.
Core Todo.txt Format Recap
First, let’s remember the flexible structure of a valid todo.txt entry:
- Optional completion marker
x(at the start) - Optional priority (e.g.,
(A)for highest,(Z)for lowest) - Optional completion date (YYYY-MM-DD, only if marked as complete)
- Mandatory creation date (YYYY-MM-DD)
- Task description
- Optional contexts (prefixed with
@, e.g.,@work) - Optional projects (prefixed with
+, e.g.,+home-improvement)
The Parser Function
We’ll use regular expressions to reliably extract each component—this handles inconsistent spacing and optional fields better than manual string splitting.
import re from typing import Dict, Optional, List def parse_todo(todo_str: str) -> Dict[str, Optional[str] | List[str] | bool]: # Initialize default values for all fields to avoid missing keys todo_dict = { "completed": False, "priority": None, "completion_date": None, "creation_date": None, "description": "", "contexts": [], "projects": [] } # Trim leading/trailing whitespace first to clean up messy input cleaned_str = todo_str.strip() if not cleaned_str: return todo_dict # Regex pattern to match all optional leading components and capture key parts # Groups: 1=completed marker, 3=priority, 4=completion date, 5=creation date, 6=remaining content pattern = r"^(\s*x\s*)?(\(([A-Z])\)\s*)?(\d{4}-\d{2}-\d{2}\s*)?(\d{4}-\d{2}-\d{2}\s*)?(.*)$" match = re.match(pattern, cleaned_str, re.IGNORECASE) if match: # Mark task as completed if the 'x' marker exists if match.group(1): todo_dict["completed"] = True # Extract priority (enforce uppercase per todo.txt spec) priority = match.group(3) if priority: todo_dict["priority"] = priority.upper() # Handle date mapping: completion date only exists if task is marked complete completion_date = match.group(4) creation_date = match.group(5) if todo_dict["completed"]: if completion_date: todo_dict["completion_date"] = completion_date.strip() if creation_date: todo_dict["creation_date"] = creation_date.strip() else: # For incomplete tasks, the first date (if present) is the creation date if completion_date: todo_dict["creation_date"] = completion_date.strip() # Extract remaining content (description + contexts/projects) remaining_content = match.group(6).strip() if not remaining_content: return todo_dict # Grab all contexts (starts with @, followed by letters/numbers/hyphens) contexts = re.findall(r"@(\w[\w-]*)", remaining_content) if contexts: todo_dict["contexts"] = [ctx.strip() for ctx in contexts] # Grab all projects (starts with +, followed by letters/numbers/hyphens) projects = re.findall(r"\+(\w[\w-]*)", remaining_content) if projects: todo_dict["projects"] = [proj.strip() for proj in projects] # Clean up description by removing context/project tags cleaned_description = re.sub(r"@\w[\w-]*|\+\w[\w-]*", "", remaining_content).strip() todo_dict["description"] = cleaned_description return todo_dict
How It Works (Key Breakdown)
Let’s walk through the critical parts:
- Default Dictionary: We start with all fields initialized to avoid missing keys, even for incomplete entries.
- Regex Pattern: The pattern accounts for optional leading components and captures each part into distinct groups, handling messy spacing gracefully.
- Date Logic: We check if the task is completed to correctly map dates—completion dates only exist if the task has the
xmarker. - Contexts/Projects: We use
re.findallto grab all instances of@contextand+project, then clean up the description by stripping these tags. - Edge Cases: Handles empty strings, entries without priority/dates, and inconsistent spacing between components.
Test It Out
Let’s run a few examples to verify the function works as expected:
# Example 1: Complete task with priority, dates, context, and project todo1 = "x (A) 2024-05-20 2024-05-15 Finish writing parser @work +coding-project" print(parse_todo(todo1)) # Output: # { # 'completed': True, # 'priority': 'A', # 'completion_date': '2024-05-20', # 'creation_date': '2024-05-15', # 'description': 'Finish writing parser', # 'contexts': ['work'], # 'projects': ['coding-project'] # } # Example 2: Incomplete task with no priority, only creation date todo2 = "2024-05-16 Buy groceries @errands" print(parse_todo(todo2)) # Output: # { # 'completed': False, # 'priority': None, # 'completion_date': None, # 'creation_date': '2024-05-16', # 'description': 'Buy groceries', # 'contexts': ['errands'], # 'projects': [] # } # Example 3: Minimal task with no extra fields todo3 = "Call mom" print(parse_todo(todo3)) # Output: # { # 'completed': False, # 'priority': None, # 'completion_date': None, # 'creation_date': None, # 'description': 'Call mom', # 'contexts': [], # 'projects': [] # }
This function strictly adheres to the todo.txt spec and handles all common (and edge) cases you’ll encounter.
内容的提问来源于stack exchange,提问作者defsoul
相关产品推荐
相关产品推荐

