在R中从corpus查找已删除文档及mydat文本预处理问询
Hey there! Let's tackle your two R-related tasks step by step, using your provided text data as an example.
Whether you're using the tm or quanteda package (the most common for corpus handling in R), here's how you can spot deleted or empty documents:
Using the tm Package
Deleted documents often either have a metadata tag marking them as removed, or their content is empty.
library(tm) # Assume your corpus is named `my_corpus` # 1. Check metadata for deletion flags print(meta(my_corpus)) # Look for fields like "deleted" or "status" deleted_docs <- my_corpus[meta(my_corpus, tag = "deleted") == TRUE] # 2. Fallback: Check for empty document content doc_content <- sapply(my_corpus, content) empty_doc_indices <- which(doc_content == "" | is.na(doc_content)) deleted_docs <- my_corpus[empty_doc_indices]
Using the quanteda Package
Quanteda uses document variables and text accessors to check for missing/deleted content:
library(quanteda) # Assume your corpus is named `my_corpus` # 1. Check document variables for deletion markers print(docvars(my_corpus)) deleted_docs <- corpus_subset(my_corpus, docvars(my_corpus, "deleted") == TRUE) # 2. Check for empty or missing text empty_docs <- corpus_subset(my_corpus, is.na(texts(my_corpus)) | texts(my_corpus) == "")
mydat Your provided text contains business operations with structured details (specs, dates, locations, etc.). We'll turn this unstructured text into a clean, analyzable format.
First, let's load the raw text into mydat:
mydat <- data.frame( raw_text = c("制作1.2x2规格横幅、切割制作2330*600mm规格板材、配送;2014年4月在奥尔忠尼启则大街(TSUM)-格尔岑侧A2投放0.85*0.65规格海报广告;制作3.7×2.7规格横幅;2011年12月1日-14日在奥尔忠尼启则大街60号A面投放3*4规格prismatron广告;在马利...投放3*12规格multipanel广告。") )
Step 1: Split into Individual Entries
The text uses semicolons to separate distinct operations—let's split them into separate rows:
library(dplyr) library(stringr) library(tidytext) mydat <- mydat %>% mutate(entry = str_split(raw_text, ";|;")) %>% unnest(entry) %>% mutate(entry = trimws(entry)) # Remove leading/trailing whitespace
Step 2: Extract Structured Key Information
We'll use regex to pull out critical details like operation type, specs, dates, locations, and item types:
# Extract operation type (制作/切割/配送/投放) mydat <- mydat %>% mutate(operation = str_extract(entry, "^制作|切割|配送|投放")) # Extract specs (matches formats like 1.2x2, 2330*600mm, 3.7×2.7) mydat <- mydat %>% mutate(spec = str_extract(entry, "\\d+[\\.x×*]\\d+[a-zA-Zmm]*")) # Extract time (matches dates like 2014年4月, 2011年12月1日-14日) mydat <- mydat %>% mutate(time = str_extract(entry, "\\d{4}年\\d{1,2}月(\\d{1,2}日-\\d{1,2}日)?")) # Extract location (pulls text between "在" and the next key word like "投放"/"制作") mydat <- mydat %>% mutate(location = str_extract(entry, "在[\u4e00-\u9fa5\\(\\)A-Z0-9-]+(?=投放|制作)")) %>% mutate(location = str_remove(location, "^在")) # Remove leading "在" # Extract item type (横幅/板材/海报广告/etc.) mydat <- mydat %>% mutate(item_type = str_extract(entry, "横幅|板材|海报广告|prismatron广告|multipanel广告"))
Step 3: Clean Up Remaining Noise
# Remove ellipses and extra symbols mydat <- mydat %>% mutate(entry = str_remove_all(entry, "\\.{2,}")) # Mark incomplete locations (like "马利...") as needing completion mydat <- mydat %>% mutate(location = ifelse(is.na(location) & operation == "投放", "待补全", location))
After these steps, mydat will be a structured data frame ready for further analysis (like counting operation types, tracking spec frequencies, or analyzing temporal trends).
内容的提问来源于stack exchange,提问作者psysky

