Python匹配问题:从pickle文件提取匹配词返回空列表求助
Hey there, let's walk through the possible issues with your code that's returning an empty list, and fix them step by step.
First, let's clean up your code for clarity (I've added missing imports since they're likely omitted in your snippet):
import pickle import re def fetch_data(document): with open('data_file.pickle', 'rb') as fp: datafile = pickle.load(fp) matched_word = [] for data in datafile.splitlines(): job_regex = r'[^a-zA-Z]' + data + r'[^a-zA-Z]' regular_expression = re.compile(job_regex, re.IGNORECASE) regex_result = re.search(regular_expression, document) if regex_result: matched_word.append(data) return matched_word
1. Pickle File Content Type Mismatch
The first critical issue is how you're handling datafile:
- You call
splitlines()on it, which assumesdatafileis a string. But if you saved a list of keywords directly to the pickle file (instead of a newline-separated string),splitlines()will either throw an error or return unexpected empty results. - Add a quick debug check right after loading the pickle file to verify:
If it's already a list, skip theprint(type(datafile)) print(datafile)splitlines()call and loop directly overdatafileinstead.
2. Overly Strict Regular Expression
Your regex pattern is the most likely culprit for missing matches:
- The
[^a-zA-Z]requires non-alphabet characters immediately before and after your keyword. This means:- Matches at the start of
document(e.g.,document = "Python developer"anddata = "Python") will fail—there's no character before "Python" to match the regex. - Matches at the end of
document(e.g.,document = "I love Python"anddata = "Python") will also fail.
- Matches at the start of
- If your keywords contain regex special characters (like
.,*,?), they'll be interpreted as regex syntax instead of literal characters. For example, a keyword likeData.Scientistwould matchDataXScientistinstead of only literal dots.
Fix the Regex:
Use word boundaries (\b) to match whole words, and escape the keyword to handle special characters:
escaped_data = re.escape(data) # Word boundaries work for matches at start/end of text, or around punctuation job_regex = r'\b' + escaped_data + r'\b' regular_expression = re.compile(job_regex, re.IGNORECASE)
3. Empty or Malformed Pickle Content
- If
datafile.splitlines()returns an empty list (because the pickle file's string content is blank), the loop won't run at all, giving you an empty result. Double-check that your pickle file actually contains the keywords you expect. - If keywords have extra whitespace or line breaks within them,
splitlines()will split them into invalid fragments that can't match your document.
4. Edge Case: Blank Keyword Entries
If your pickle file has blank lines, splitlines() will include empty strings in your loop. These will never match anything, but you can skip them to avoid unnecessary checks:
for data in datafile.splitlines(): cleaned_data = data.strip() if not cleaned_data: continue # Rest of your regex logic using cleaned_data
Final Fixed Code Example
Here's the revised code incorporating all these fixes:
import pickle import re def fetch_data(document): with open('data_file.pickle', 'rb') as fp: datafile = pickle.load(fp) matched_word = [] # Handle both string and list inputs from pickle if isinstance(datafile, str): keywords = [line.strip() for line in datafile.splitlines() if line.strip()] else: keywords = [str(k).strip() for k in datafile if str(k).strip()] for keyword in keywords: escaped_keyword = re.escape(keyword) regex_pattern = re.compile(r'\b' + escaped_keyword + r'\b', re.IGNORECASE) if regex_pattern.search(document): matched_word.append(keyword) return matched_word
内容的提问来源于stack exchange,提问作者Jay Pratap Pandey

