You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于Python 3实现todo.txt:从字符串提取待办任务信息至字典

Python 3 Todo.txt Parser Function

Got it, let's build a robust function to parse todo.txt entries into a structured dictionary. I’ve worked with this format extensively, so I’ll break down each component step by step to handle all edge cases.

Core Todo.txt Format Recap

First, let’s remember the flexible structure of a valid todo.txt entry:

  • Optional completion marker x (at the start)
  • Optional priority (e.g., (A) for highest, (Z) for lowest)
  • Optional completion date (YYYY-MM-DD, only if marked as complete)
  • Mandatory creation date (YYYY-MM-DD)
  • Task description
  • Optional contexts (prefixed with @, e.g., @work)
  • Optional projects (prefixed with +, e.g., +home-improvement)

The Parser Function

We’ll use regular expressions to reliably extract each component—this handles inconsistent spacing and optional fields better than manual string splitting.

import re
from typing import Dict, Optional, List

def parse_todo(todo_str: str) -> Dict[str, Optional[str] | List[str] | bool]:
    # Initialize default values for all fields to avoid missing keys
    todo_dict = {
        "completed": False,
        "priority": None,
        "completion_date": None,
        "creation_date": None,
        "description": "",
        "contexts": [],
        "projects": []
    }
    
    # Trim leading/trailing whitespace first to clean up messy input
    cleaned_str = todo_str.strip()
    if not cleaned_str:
        return todo_dict
    
    # Regex pattern to match all optional leading components and capture key parts
    # Groups: 1=completed marker, 3=priority, 4=completion date, 5=creation date, 6=remaining content
    pattern = r"^(\s*x\s*)?(\(([A-Z])\)\s*)?(\d{4}-\d{2}-\d{2}\s*)?(\d{4}-\d{2}-\d{2}\s*)?(.*)$"
    match = re.match(pattern, cleaned_str, re.IGNORECASE)
    
    if match:
        # Mark task as completed if the 'x' marker exists
        if match.group(1):
            todo_dict["completed"] = True
        
        # Extract priority (enforce uppercase per todo.txt spec)
        priority = match.group(3)
        if priority:
            todo_dict["priority"] = priority.upper()
        
        # Handle date mapping: completion date only exists if task is marked complete
        completion_date = match.group(4)
        creation_date = match.group(5)
        
        if todo_dict["completed"]:
            if completion_date:
                todo_dict["completion_date"] = completion_date.strip()
            if creation_date:
                todo_dict["creation_date"] = creation_date.strip()
        else:
            # For incomplete tasks, the first date (if present) is the creation date
            if completion_date:
                todo_dict["creation_date"] = completion_date.strip()
        
        # Extract remaining content (description + contexts/projects)
        remaining_content = match.group(6).strip()
        if not remaining_content:
            return todo_dict
        
        # Grab all contexts (starts with @, followed by letters/numbers/hyphens)
        contexts = re.findall(r"@(\w[\w-]*)", remaining_content)
        if contexts:
            todo_dict["contexts"] = [ctx.strip() for ctx in contexts]
        
        # Grab all projects (starts with +, followed by letters/numbers/hyphens)
        projects = re.findall(r"\+(\w[\w-]*)", remaining_content)
        if projects:
            todo_dict["projects"] = [proj.strip() for proj in projects]
        
        # Clean up description by removing context/project tags
        cleaned_description = re.sub(r"@\w[\w-]*|\+\w[\w-]*", "", remaining_content).strip()
        todo_dict["description"] = cleaned_description
    
    return todo_dict

How It Works (Key Breakdown)

Let’s walk through the critical parts:

  • Default Dictionary: We start with all fields initialized to avoid missing keys, even for incomplete entries.
  • Regex Pattern: The pattern accounts for optional leading components and captures each part into distinct groups, handling messy spacing gracefully.
  • Date Logic: We check if the task is completed to correctly map dates—completion dates only exist if the task has the x marker.
  • Contexts/Projects: We use re.findall to grab all instances of @context and +project, then clean up the description by stripping these tags.
  • Edge Cases: Handles empty strings, entries without priority/dates, and inconsistent spacing between components.

Test It Out

Let’s run a few examples to verify the function works as expected:

# Example 1: Complete task with priority, dates, context, and project
todo1 = "x (A) 2024-05-20 2024-05-15 Finish writing parser @work +coding-project"
print(parse_todo(todo1))
# Output:
# {
#     'completed': True,
#     'priority': 'A',
#     'completion_date': '2024-05-20',
#     'creation_date': '2024-05-15',
#     'description': 'Finish writing parser',
#     'contexts': ['work'],
#     'projects': ['coding-project']
# }

# Example 2: Incomplete task with no priority, only creation date
todo2 = "2024-05-16 Buy groceries @errands"
print(parse_todo(todo2))
# Output:
# {
#     'completed': False,
#     'priority': None,
#     'completion_date': None,
#     'creation_date': '2024-05-16',
#     'description': 'Buy groceries',
#     'contexts': ['errands'],
#     'projects': []
# }

# Example 3: Minimal task with no extra fields
todo3 = "Call mom"
print(parse_todo(todo3))
# Output:
# {
#     'completed': False,
#     'priority': None,
#     'completion_date': None,
#     'creation_date': None,
#     'description': 'Call mom',
#     'contexts': [],
#     'projects': []
# }

This function strictly adheres to the todo.txt spec and handles all common (and edge) cases you’ll encounter.

内容的提问来源于stack exchange,提问作者defsoul

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 08:23:28