You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python 3中通过for循环提取正则匹配结果存入字典并去除None值的问题排查

Fixing Your Regex Matching & Dictionary Cleanup Issue

Let's break down what's going wrong with your code and fix it to meet your goal of only keeping valid, non-None matches in your dictionaries.

Key Problems in Your Current Code

  1. You're iterating over words, not lines
    Your code uses fh.split() which splits the entire file content by whitespace (spaces, tabs, newlines) into individual words. That's why you're getting so many dictionaries with None values—most words don't match your date/time patterns. You need to iterate over lines instead.

  2. You're forcing every dictionary to have date and time keys
    Even when a match fails, you're adding the key with a None value. To keep only valid matches, you should only add the key to the dictionary if the regex actually finds a match.

Corrected Code

import re
import pandas as pd

list1 = []

# Use a with-statement to safely handle file I/O (auto-closes the file)
with open(r"test_data.txt", "r") as fh:
    # Iterate over each line in the file directly
    for line in fh:
        stripped_line = line.strip()  # Remove leading/trailing whitespace and newlines
        if not stripped_line:
            continue  # Skip empty lines to avoid empty dictionaries
        
        list_dict = {}
        
        # Match date pattern only add key if found
        date_match = re.search(r"(\d{1})[/.-](\d{1})[/.-](\d{4})$", stripped_line)
        if date_match:
            list_dict["date"] = date_match.group()
        
        # Match time pattern only add key if found
        time_match = re.search(r"(\d{1,2})[:](\d{2})[:](\d{2})$", stripped_line)
        if time_match:
            list_dict["time"] = time_match.group()
        
        # Only add the dictionary to the list if it has at least one valid field
        if list_dict:
            list1.append(list_dict)

print(list1)
df = pd.DataFrame(list1)
df.to_csv("test_export_clean.csv", index=False)

How This Differs From the Reference Code

The DataQuest code splits the input into email blocks (each block is a full email), then extracts multiple fields from each block—even if some fields are None (since emails typically have senders, recipients, etc., even if parsing fails). Your use case is different: you want one dictionary per line, and only keep keys for fields that actually match.

Optimization Tips for Python Beginners

  1. Precompile regex patterns
    For large files, compiling your regex once before looping will speed up processing:

    # Define patterns once at the top
    date_pattern = re.compile(r"(\d{1})[/.-](\d{1})[/.-](\d{4})$")
    time_pattern = re.compile(r"(\d{1,2})[:](\d{2})[:](\d{2})$")
    
    # Then use them in the loop:
    date_match = date_pattern.search(stripped_line)
    
  2. Encapsulate extraction logic in a function
    Makes your code cleaner and easier to maintain:

    def extract_date_time(line):
        fields = {}
        date_match = date_pattern.search(line)
        if date_match:
            fields["date"] = date_match.group()
        time_match = time_pattern.search(line)
        if time_match:
            fields["time"] = time_match.group()
        return fields
    
    # In the loop:
    line_fields = extract_date_time(stripped_line)
    if line_fields:
        list1.append(line_fields)
    
  3. Handle file encoding explicitly
    If your text file uses non-ASCII characters, add an encoding parameter to open():

    with open(r"test_data.txt", "r", encoding="utf-8") as fh:
    

内容的提问来源于stack exchange,提问作者dpec

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.29 11:12:44