Python 3中通过for循环提取正则匹配结果存入字典并去除None值的问题排查
Let's break down what's going wrong with your code and fix it to meet your goal of only keeping valid, non-None matches in your dictionaries.
Key Problems in Your Current Code
You're iterating over words, not lines
Your code usesfh.split()which splits the entire file content by whitespace (spaces, tabs, newlines) into individual words. That's why you're getting so many dictionaries withNonevalues—most words don't match your date/time patterns. You need to iterate over lines instead.You're forcing every dictionary to have
dateandtimekeys
Even when a match fails, you're adding the key with aNonevalue. To keep only valid matches, you should only add the key to the dictionary if the regex actually finds a match.
Corrected Code
import re import pandas as pd list1 = [] # Use a with-statement to safely handle file I/O (auto-closes the file) with open(r"test_data.txt", "r") as fh: # Iterate over each line in the file directly for line in fh: stripped_line = line.strip() # Remove leading/trailing whitespace and newlines if not stripped_line: continue # Skip empty lines to avoid empty dictionaries list_dict = {} # Match date pattern only add key if found date_match = re.search(r"(\d{1})[/.-](\d{1})[/.-](\d{4})$", stripped_line) if date_match: list_dict["date"] = date_match.group() # Match time pattern only add key if found time_match = re.search(r"(\d{1,2})[:](\d{2})[:](\d{2})$", stripped_line) if time_match: list_dict["time"] = time_match.group() # Only add the dictionary to the list if it has at least one valid field if list_dict: list1.append(list_dict) print(list1) df = pd.DataFrame(list1) df.to_csv("test_export_clean.csv", index=False)
How This Differs From the Reference Code
The DataQuest code splits the input into email blocks (each block is a full email), then extracts multiple fields from each block—even if some fields are None (since emails typically have senders, recipients, etc., even if parsing fails). Your use case is different: you want one dictionary per line, and only keep keys for fields that actually match.
Optimization Tips for Python Beginners
Precompile regex patterns
For large files, compiling your regex once before looping will speed up processing:# Define patterns once at the top date_pattern = re.compile(r"(\d{1})[/.-](\d{1})[/.-](\d{4})$") time_pattern = re.compile(r"(\d{1,2})[:](\d{2})[:](\d{2})$") # Then use them in the loop: date_match = date_pattern.search(stripped_line)Encapsulate extraction logic in a function
Makes your code cleaner and easier to maintain:def extract_date_time(line): fields = {} date_match = date_pattern.search(line) if date_match: fields["date"] = date_match.group() time_match = time_pattern.search(line) if time_match: fields["time"] = time_match.group() return fields # In the loop: line_fields = extract_date_time(stripped_line) if line_fields: list1.append(line_fields)Handle file encoding explicitly
If your text file uses non-ASCII characters, add an encoding parameter toopen():with open(r"test_data.txt", "r", encoding="utf-8") as fh:
内容的提问来源于stack exchange,提问作者dpec

