You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于SpaCy与dateparser的新闻文本日期实体解析:不完整日期年份误判问题及解决方案问询

Great question—this is a classic edge case when parsing relative dates for news content, where a one-size-fits-all PREFER_DATES_FROM setting just doesn't cut it. Let's break down a solution that combines custom logic with dateparser to fix both your examples.

Core Idea

Instead of forcing dateparser to always prefer past or future dates, we'll:

  1. Distinguish between dates with explicit years and those without (since explicit years don't need guessing)
  2. Check for contextual keywords (like "next" or "last") that explicitly signal past/future
  3. Use time-difference logic for ambiguous, year-less dates to decide if they refer to the same year or the next

Custom Date Parsing Function

Here's a Python 3.8-compatible function that implements this logic:

import re
from dateparser import parse
from datetime import datetime, timedelta

def parse_news_date(date_str, base_date):
    # Check if the date string contains an explicit year
    has_year = bool(re.search(r'\b(19|20)\d{2}\b', date_str.lower()))
    
    # Define keywords that explicitly signal past/future context
    future_triggers = {"next", "upcoming", "future", "will", "planned", "scheduled", "coming"}
    past_triggers = {"last", "previous", "past", "already", "announced", "released"}
    
    has_future_context = any(word in date_str.lower() for word in future_triggers)
    has_past_context = any(word in date_str.lower() for word in past_triggers)

    if has_year:
        # Explicit year: parse normally using the article's publish date as base
        return parse(
            date_str,
            languages=['en'],
            settings={
                'RELATIVE_BASE': base_date,
                'PREFER_DAY_OF_MONTH': 'last'
            }
        )
    elif has_future_context:
        # Clear future signal: force dateparser to prefer future dates
        return parse(
            date_str,
            languages=['en'],
            settings={
                'RELATIVE_BASE': base_date,
                'PREFER_DAY_OF_MONTH': 'last',
                'PREFER_DATES_FROM': 'future'
            }
        )
    elif has_past_context:
        # Clear past signal: force dateparser to prefer past dates
        return parse(
            date_str,
            languages=['en'],
            settings={
                'RELATIVE_BASE': base_date,
                'PREFER_DAY_OF_MONTH': 'last',
                'PREFER_DATES_FROM': 'past'
            }
        )
    else:
        # Ambiguous year-less date: use time-difference logic
        # First parse as same year (force past to avoid jumping to next year)
        parsed_same_year = parse(
            date_str,
            languages=['en'],
            settings={
                'RELATIVE_BASE': base_date,
                'PREFER_DAY_OF_MONTH': 'last',
                'PREFER_DATES_FROM': 'past'
            }
        )
        
        if parsed_same_year is None:
            return None  # Handle unparseable dates gracefully
        
        # Calculate days between parsed date and article publish date
        days_since_parsed = (base_date - parsed_same_year).days
        
        # If the parsed date is more than 6 months in the past, assume it refers to next year
        # (News rarely references events that far in the past without context)
        if days_since_parsed > 180:
            return parsed_same_year.replace(year=parsed_same_year.year + 1)
        else:
            return parsed_same_year

Integrate with Your Existing Code

Update your loop to use this function, and pre-convert your publish dates to datetime objects for efficiency:

# Pre-convert publish dates to datetime to avoid repeated parsing in the loop
df_test['RP_Date_Datetime'] = df_test['RP_DateFormatted'].apply(lambda x: datetime.strptime(x, '%Y-%m-%d'))

for index, row in df_test.iterrows():
    doc = nlp(row.Text_4)
    entities = {key: list(g) for key, g in groupby(sorted(doc.ents, key=lambda x: x.label_), lambda x: x.label_)}
    
    if 'DATE' in entities:
        # Use our custom function to parse each date entity
        parsed_dates = [
            parse_news_date(ent.text, row['RP_Date_Datetime']) 
            for ent in entities['DATE']
        ]
        # Use .at to avoid SettingWithCopyWarning
        df_test.at[index, 'PY_Entities_DatesParsed'] = parsed_dates
    else:
        df_test.at[index, 'PY_Entities_DatesParsed'] = []

Key Adjustments for Your Use Cases

  • Your first example (publish date 2005-08-15, date string "Aug 15"): The function parses it as 2005-08-15 (same year) since the time difference is 0 days (well under 180).
  • Your second example (publish date 2005-08-15, date string "Feb"): The function parses it as 2005-02-01 first, calculates a 195-day difference (over 180), and adjusts it to 2006-02-01 (correctly identifying it as a future event).

Optional Tweaks

  • Adjust the 180-day threshold based on your news type (e.g., use 90 days for tech news where product launches are planned months in advance).
  • Expand the future_triggers/past_triggers sets to include domain-specific terms (like "launching" or "unveiled").
  • Add error handling for dateparser returning None (e.g., log unparseable dates for manual review).

内容的提问来源于stack exchange,提问作者AlexanderP

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.01 00:52:35