You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在R中从corpus查找已删除文档及mydat文本预处理问询

Hey there! Let's tackle your two R-related tasks step by step, using your provided text data as an example.

1. Detecting Deleted Documents in an R Corpus

Whether you're using the tm or quanteda package (the most common for corpus handling in R), here's how you can spot deleted or empty documents:

Using the tm Package

Deleted documents often either have a metadata tag marking them as removed, or their content is empty.

library(tm)

# Assume your corpus is named `my_corpus`
# 1. Check metadata for deletion flags
print(meta(my_corpus)) # Look for fields like "deleted" or "status"
deleted_docs <- my_corpus[meta(my_corpus, tag = "deleted") == TRUE]

# 2. Fallback: Check for empty document content
doc_content <- sapply(my_corpus, content)
empty_doc_indices <- which(doc_content == "" | is.na(doc_content))
deleted_docs <- my_corpus[empty_doc_indices]

Using the quanteda Package

Quanteda uses document variables and text accessors to check for missing/deleted content:

library(quanteda)

# Assume your corpus is named `my_corpus`
# 1. Check document variables for deletion markers
print(docvars(my_corpus))
deleted_docs <- corpus_subset(my_corpus, docvars(my_corpus, "deleted") == TRUE)

# 2. Check for empty or missing text
empty_docs <- corpus_subset(my_corpus, is.na(texts(my_corpus)) | texts(my_corpus) == "")
2. Text Preprocessing for mydat

Your provided text contains business operations with structured details (specs, dates, locations, etc.). We'll turn this unstructured text into a clean, analyzable format.

First, let's load the raw text into mydat:

mydat <- data.frame(
  raw_text = c("制作1.2x2规格横幅、切割制作2330*600mm规格板材、配送;2014年4月在奥尔忠尼启则大街(TSUM)-格尔岑侧A2投放0.85*0.65规格海报广告;制作3.7×2.7规格横幅;2011年12月1日-14日在奥尔忠尼启则大街60号A面投放3*4规格prismatron广告;在马利...投放3*12规格multipanel广告。")
)

Step 1: Split into Individual Entries

The text uses semicolons to separate distinct operations—let's split them into separate rows:

library(dplyr)
library(stringr)
library(tidytext)

mydat <- mydat %>%
  mutate(entry = str_split(raw_text, ";|;")) %>%
  unnest(entry) %>%
  mutate(entry = trimws(entry)) # Remove leading/trailing whitespace

Step 2: Extract Structured Key Information

We'll use regex to pull out critical details like operation type, specs, dates, locations, and item types:

# Extract operation type (制作/切割/配送/投放)
mydat <- mydat %>%
  mutate(operation = str_extract(entry, "^制作|切割|配送|投放"))

# Extract specs (matches formats like 1.2x2, 2330*600mm, 3.7×2.7)
mydat <- mydat %>%
  mutate(spec = str_extract(entry, "\\d+[\\.x×*]\\d+[a-zA-Zmm]*"))

# Extract time (matches dates like 2014年4月, 2011年12月1日-14日)
mydat <- mydat %>%
  mutate(time = str_extract(entry, "\\d{4}年\\d{1,2}月(\\d{1,2}日-\\d{1,2}日)?"))

# Extract location (pulls text between "在" and the next key word like "投放"/"制作")
mydat <- mydat %>%
  mutate(location = str_extract(entry, "在[\u4e00-\u9fa5\\(\\)A-Z0-9-]+(?=投放|制作)")) %>%
  mutate(location = str_remove(location, "^在")) # Remove leading "在"

# Extract item type (横幅/板材/海报广告/etc.)
mydat <- mydat %>%
  mutate(item_type = str_extract(entry, "横幅|板材|海报广告|prismatron广告|multipanel广告"))

Step 3: Clean Up Remaining Noise

# Remove ellipses and extra symbols
mydat <- mydat %>%
  mutate(entry = str_remove_all(entry, "\\.{2,}"))

# Mark incomplete locations (like "马利...") as needing completion
mydat <- mydat %>%
  mutate(location = ifelse(is.na(location) & operation == "投放", "待补全", location))

After these steps, mydat will be a structured data frame ready for further analysis (like counting operation types, tracking spec frequencies, or analyzing temporal trends).

内容的提问来源于stack exchange,提问作者psysky

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.20 10:39:00