You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python脚本IndexError: list index out of range及datefinder异常求助

Fixing Your IndexError & Date Extraction Issues

Let's tackle the problems you're encountering one by one, with practical fixes and improvements:

1. Why You're Getting IndexError: list index out of range

Your code assumes datefinder.find_dates() returns at least 2 dates whenever it finds any, so you directly access matches[1]. But in reality:

  • Some files might only have 1 valid date (like just the article's publication date)
  • datefinder might miss dates entirely (you handle this case, but not the 1-date scenario)

Quick Fix: Check the length of matches first, and pick the most relevant date (usually the first one for publication dates):

matches = list(datefinder.find_dates(text1))
if len(matches) > 0:
    # Use the first detected date (adjust index if you need a different one)
    date = matches[0]
    strdate = str(date)
else:
    print(f'No dates found in {filename}')  # Fixed Python 3 print syntax
    strdate = ''

2. Fixing Missing/Wrong Dates from datefinder

datefinder is handy but can miss edge-case mm-dd-yyyy formats (especially with time). Since you know the exact date format you need, add a regex fallback to explicitly capture these:

import re
from datetime import datetime

# Add this inside extract_data(), before datefinder logic
date_regex = r'\b(\d{2}-\d{2}-\d{4}(?: \d{1,2}:\d{2})?)\b'
regex_matches = re.findall(date_regex, text1)

# Prioritize regex results (since you know the exact format)
if regex_matches:
    try:
        # Parse with time first; fall back to date-only if needed
        date = datetime.strptime(regex_matches[0], '%m-%d-%Y %H:%M')
    except ValueError:
        date = datetime.strptime(regex_matches[0], '%m-%d-%Y')
    strdate = str(date)
else:
    # Fall back to datefinder
    matches = list(datefinder.find_dates(text1))
    if len(matches) > 0:
        date = matches[0]
        strdate = str(date)
    else:
        print(f'No dates found in {filename}')
        strdate = ''

3. Other Critical Fixes for Your Script

  • Python 3 Compatibility: Your original print statement uses Python 2 syntax—add parentheses or use f-strings for readability.
  • Unsafe re.search Call: If a file doesn't contain "X words", re.search(r'(.*) words', text1) returns None, and calling .group(1) will crash. Add a safety check:
    count_match = re.search(r'(.*) words', text1)
    matchcount = count_match.group(1).strip() if count_match else '0'  # Handle missing data
    
  • Path Escaping: Use raw strings for Windows paths to avoid unintended escape sequences (like \t being treated as a tab):
    os.chdir(r'C:\Users\dul\Dropbox\Article\parsedarticles')
    files = os.listdir(r'C:\Users\dul\Dropbox\Article\parsedarticles')
    

Modified Full Script

Here's the updated code with all fixes included:

import os, datefinder, re
from datetime import datetime

# Use raw string for Windows path to avoid escape issues
os.chdir(r'C:\Users\dul\Dropbox\Article\parsedarticles')

def matchwho(text_to_match):
    if 'This story was generated by' in text_to_match:
        return('1')
    elif any(phrase in text_to_match for phrase in [
        'This story includes elements generated',
        'Elements of this story were generated',
        'Portions of this story were generated',
        'Parts of this story were generated',
        'A portion of this story was generated',
        'This sory was partially generated by',  # Note: typo "sory" -> "story"?
        'This story contains elements generated by',
        'This story includes information generated by'
    ]):
        return('2')
    elif 'This story was originally generated by' in text_to_match:
        return('1')
    else:
        return('3')

def extract_data(filename):
    with open(filename, 'r') as file1:
        text1 = file1.read()
    
    # Extract date: regex first for mm-dd-yyyy, then datefinder fallback
    date_regex = r'\b(\d{2}-\d{2}-\d{4}(?: \d{1,2}:\d{2})?)\b'
    regex_matches = re.findall(date_regex, text1)
    strdate = ''
    if regex_matches:
        try:
            date = datetime.strptime(regex_matches[0], '%m-%d-%Y %H:%M')
        except ValueError:
            date = datetime.strptime(regex_matches[0], '%m-%d-%Y')
        strdate = str(date)
    else:
        matches = list(datefinder.find_dates(text1))
        if len(matches) > 0:
            date = matches[0]
            strdate = str(date)
        else:
            print(f'No dates found in {filename}')
    
    # Extract company name (note: current logic is fragile—adjust if needed)
    matchcomp2 = text1.split(' ', 1)[0] if text1.strip() else ''
    
    # Extract word count safely
    count_match = re.search(r'(.*) words', text1)
    matchcount = count_match.group(1).strip() if count_match else '0'
    
    # Determine article type
    matchw = str(matchwho(text1))
    
    # Combine results
    combid = f"{matchcomp2},{strdate},{matchw},{matchcount}"
    
    # Append to output file
    with open('outfile.txt', "a+") as outfile:
        outfile.write(f"\n{combid}")

# Process all .txt files
files = os.listdir(r'C:\Users\dul\Dropbox\Article\parsedarticles')
for file in files:
    if file.endswith(".txt"):
        extract_data(file)

Quick Notes

  • I cleaned up the matchwho function to use any() for better readability.
  • There's a typo in your original code: This sory was partially generated by (should be "story")—I left it as-is but flagged it.
  • The company name extraction (text1.split(' ',1)[0]) is very fragile—if the first word isn't the company name, you'll need to adjust this logic (e.g., use regex to target company names specifically).

内容的提问来源于stack exchange,提问作者Philip

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.12 03:51:23