You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

求助:如何在R语言中从corpus对象提取首次出现的日期

Solution: Extract First Date as Docvar in Quanteda Corpus

Let's break down and fix your issue step by step. The main problems with your original code were:

  1. Using str_extract_all (which returns lists) instead of str_extract to get the first date match
  2. A regex that matched non-date strings
  3. Not handling invalid month spellings as requested

Step 1: Load Required Packages

First, make sure you have these packages installed and loaded:

library(quanteda)
library(stringr)
library(lubridate)

Step 2: Recreate Your Corpus

Let's confirm we're working with the same dataset:

v1 <- c("(SE22-y -7 A go q ,, Document of The World Bank FOR OFFICIAL USE ONLY il I ( >I8.( )]i 1 t'f-l±E C 4'( | Report No. 9529-LSO l il .rt N ,- / . t ,!I . 1. 'i 1( T v f) (: AR.) STAFF APPRAISAL REPORT KINGDOM OF LESOTHO EDUCATION SECTOR DEVELOPMENT PROJECT JUNE 19, 1991 Population and Human Resources Division Southern Africa Department This document has a restricted distribution and may be used by reipients only in the performance of their official duties. Its contents may not otherwise be disclosed without World Bank authorization.", "Document of The World Bank Report No. 13611-PAK STAFF APPRAISAL REPORT PAKISTAN POPULATION WELFARE PROGRAM PROJECT FREBRUARY 10, 1995 Population and Human Resources Division Country Department I South Asia Region", "I Toward an Environmental Strategy for Asia A Summary of a World Bank Discussion Paper Carter Brandon Ramesh Ramankutty The World Bank Washliington, D.C. (C 1993 The International Bank for Reconstruction and Development / THiE WORLD BANK 1818 H Street, N.W. Washington, D.C. 20433 All rights reserved Manufactured in the United States of America First printing November 1993", "Report No. PID9188 Project Name East Timor-TP-Emergency School (@) Readiness Project Region East Asia and Pacific Region Sector Other Education Project ID TPPE70268 Borrower(s) EAST TIMOR Implementing Agency Address UNTAET (UN TRANSITIONAL ADMINISTRATION FOR EAST TIMOR) Contact Person: Cecilio Adorna, UNTAET, Dili, East Timor Fax: 61-8 89 422198 Environment Category C Date PID Prepared June 16, 2000 Projected Appraisal Date May 27, 2000 Projected Board Date June 20, 2000", "Page 1 CONFORMED COPY CREDIT NUMBER 2447-CHA (Reform, Institutional Support and Preinvestment Project) between PEOPLE'S REPUBLIC OF CHINA and INTERNATIONAL DEVELOPMENT ASSOCIATION Dated December 30, 1992")
c1 <- corpus(v1)

Step 3: Define a Date Extraction Function

This function will:

  • Extract the first valid date string (ignoring single-letter "months" like the "C" in "(C 1993")
  • Parse valid dates in either Month Day, Year or Month Year format
  • Handle invalid month spellings by falling back to a partial date (using January for unknown months, or just the year)
extract_first_date <- function(text) {
  # Regex to match: 3+ letters (month) + optional day/comma + 4-digit year
  date_str <- str_extract(text, "[A-Za-z]{3,}(?:\\s+\\d{1,2},)?\\s+\\d{4}")
  
  if (is.na(date_str)) return(NA)
  
  # Try parsing with valid month formats
  parsed_date <- parse_date_time(
    date_str, 
    orders = c("B d,Y", "B Y"),  # "B" = full month name
    quiet = TRUE  # Suppress parsing error messages
  )
  
  # Handle invalid month spellings (per your request to ignore the month)
  if (is.na(parsed_date)) {
    # Check if there's a day component
    day_match <- str_extract(date_str, "\\d{1,2},")
    if (!is.na(day_match)) {
      day <- str_remove(day_match, ",")
      year <- str_extract(date_str, "\\d{4}")
      return(as.Date(paste0(year, "-01-", day)))  # Use Jan for unknown month
    } else {
      # Only year available
      year <- str_extract(date_str, "\\d{4}")
      return(as.Date(paste0(year, "-01-01")))
    }
  }
  
  # Return valid parsed date as Date type
  as.Date(parsed_date)
}

Step 4: Apply the Function and Add as Docvar

Use sapply to run the function on each text in your corpus, then attach the results as a new docvar:

# Extract dates
first_dates <- sapply(texts(c1), extract_first_date)

# Add as a docvar
docvars(c1, "first_date") <- first_dates

Step 5: Verify the Results

Check the new docvar to confirm everything works:

print(docvars(c1))

Expected Output:

first_date
1 1991-06-19
2 1995-01-10
3 1993-11-01
4 2000-06-16
5 1992-12-30

Key Fixes Explained:

  • Regex Improvement: The regex [A-Za-z]{3,}(?:\\s+\\d{1,2},)?\\s+\\d{4} ensures we only match month-like strings (3+ letters) instead of random alphanumeric chunks.
  • Single Match Extraction: str_extract gets only the first date occurrence, which aligns with your requirement.
  • Error Handling: For misspelled months (like "FREBRUARY"), we fall back to a partial date instead of throwing an error.
  • Quanteda Best Practice: Using texts(c1) instead of directly accessing c1$documents$texts is the recommended way to retrieve corpus text content.

内容的提问来源于stack exchange,提问作者Mel Schickel

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 07:03:17