Python脚本IndexError: list index out of range及datefinder异常求助
Fixing Your IndexError & Date Extraction Issues
Let's tackle the problems you're encountering one by one, with practical fixes and improvements:
1. Why You're Getting IndexError: list index out of range
Your code assumes datefinder.find_dates() returns at least 2 dates whenever it finds any, so you directly access matches[1]. But in reality:
- Some files might only have 1 valid date (like just the article's publication date)
datefindermight miss dates entirely (you handle this case, but not the 1-date scenario)
Quick Fix: Check the length of matches first, and pick the most relevant date (usually the first one for publication dates):
matches = list(datefinder.find_dates(text1)) if len(matches) > 0: # Use the first detected date (adjust index if you need a different one) date = matches[0] strdate = str(date) else: print(f'No dates found in {filename}') # Fixed Python 3 print syntax strdate = ''
2. Fixing Missing/Wrong Dates from datefinder
datefinder is handy but can miss edge-case mm-dd-yyyy formats (especially with time). Since you know the exact date format you need, add a regex fallback to explicitly capture these:
import re from datetime import datetime # Add this inside extract_data(), before datefinder logic date_regex = r'\b(\d{2}-\d{2}-\d{4}(?: \d{1,2}:\d{2})?)\b' regex_matches = re.findall(date_regex, text1) # Prioritize regex results (since you know the exact format) if regex_matches: try: # Parse with time first; fall back to date-only if needed date = datetime.strptime(regex_matches[0], '%m-%d-%Y %H:%M') except ValueError: date = datetime.strptime(regex_matches[0], '%m-%d-%Y') strdate = str(date) else: # Fall back to datefinder matches = list(datefinder.find_dates(text1)) if len(matches) > 0: date = matches[0] strdate = str(date) else: print(f'No dates found in {filename}') strdate = ''
3. Other Critical Fixes for Your Script
- Python 3 Compatibility: Your original
printstatement uses Python 2 syntax—add parentheses or use f-strings for readability. - Unsafe
re.searchCall: If a file doesn't contain "X words",re.search(r'(.*) words', text1)returnsNone, and calling.group(1)will crash. Add a safety check:count_match = re.search(r'(.*) words', text1) matchcount = count_match.group(1).strip() if count_match else '0' # Handle missing data - Path Escaping: Use raw strings for Windows paths to avoid unintended escape sequences (like
\tbeing treated as a tab):os.chdir(r'C:\Users\dul\Dropbox\Article\parsedarticles') files = os.listdir(r'C:\Users\dul\Dropbox\Article\parsedarticles')
Modified Full Script
Here's the updated code with all fixes included:
import os, datefinder, re from datetime import datetime # Use raw string for Windows path to avoid escape issues os.chdir(r'C:\Users\dul\Dropbox\Article\parsedarticles') def matchwho(text_to_match): if 'This story was generated by' in text_to_match: return('1') elif any(phrase in text_to_match for phrase in [ 'This story includes elements generated', 'Elements of this story were generated', 'Portions of this story were generated', 'Parts of this story were generated', 'A portion of this story was generated', 'This sory was partially generated by', # Note: typo "sory" -> "story"? 'This story contains elements generated by', 'This story includes information generated by' ]): return('2') elif 'This story was originally generated by' in text_to_match: return('1') else: return('3') def extract_data(filename): with open(filename, 'r') as file1: text1 = file1.read() # Extract date: regex first for mm-dd-yyyy, then datefinder fallback date_regex = r'\b(\d{2}-\d{2}-\d{4}(?: \d{1,2}:\d{2})?)\b' regex_matches = re.findall(date_regex, text1) strdate = '' if regex_matches: try: date = datetime.strptime(regex_matches[0], '%m-%d-%Y %H:%M') except ValueError: date = datetime.strptime(regex_matches[0], '%m-%d-%Y') strdate = str(date) else: matches = list(datefinder.find_dates(text1)) if len(matches) > 0: date = matches[0] strdate = str(date) else: print(f'No dates found in {filename}') # Extract company name (note: current logic is fragile—adjust if needed) matchcomp2 = text1.split(' ', 1)[0] if text1.strip() else '' # Extract word count safely count_match = re.search(r'(.*) words', text1) matchcount = count_match.group(1).strip() if count_match else '0' # Determine article type matchw = str(matchwho(text1)) # Combine results combid = f"{matchcomp2},{strdate},{matchw},{matchcount}" # Append to output file with open('outfile.txt', "a+") as outfile: outfile.write(f"\n{combid}") # Process all .txt files files = os.listdir(r'C:\Users\dul\Dropbox\Article\parsedarticles') for file in files: if file.endswith(".txt"): extract_data(file)
Quick Notes
- I cleaned up the
matchwhofunction to useany()for better readability. - There's a typo in your original code:
This sory was partially generated by(should be "story")—I left it as-is but flagged it. - The company name extraction (
text1.split(' ',1)[0]) is very fragile—if the first word isn't the company name, you'll need to adjust this logic (e.g., use regex to target company names specifically).
内容的提问来源于stack exchange,提问作者Philip
相关产品推荐
相关产品推荐

