You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在R语言中提取结构化文本文件的PD、TD元素并存储为表格的技术求助

Extract PD and TD Elements from Text File in R

Hey there! Let's work through this problem step by step. Based on your sample text structure, each tag (like PD, TD) is a two-letter uppercase code followed by its content. We'll use regular expressions to pull out the PD (date) and TD (main text) values, then package them into a table for easy saving.

Step 1: Install & Load Required Packages

First, make sure you have the necessary tools installed. We'll use stringr for regex operations and dplyr/tibble for clean table creation:

install.packages(c("stringr", "dplyr"))

Once installed, load the packages:

library(stringr)
library(dplyr)
library(tibble)

Step 2: Read Your Text File

Read the file into a single string (this handles multi-line files too):

# Replace "your_file.txt" with your actual file path
text_content <- readLines("your_file.txt", warn = FALSE)
# Collapse any line breaks into a single continuous string
text_content <- paste(text_content, collapse = " ")

Step 3: Extract PD and TD Values

We'll use regex to target the specific tags:

  • For PD: Capture everything after "PD " until the next two-letter uppercase tag.
  • For TD: Capture everything after "TD " (since it’s likely the final tag in each entry).
# Extract PD date and remove the "PD " prefix
pd_date <- str_extract(text_content, "PD\\s+(.*?)\\s+(?=[A-Z]{2}\\s)") %>%
  str_remove("^PD\\s+")

# Extract TD content and remove the "TD " prefix
td_text <- str_extract(text_content, "TD\\s+(.*)") %>%
  str_remove("^TD\\s+")

Step 4: Create & Save the Table

Turn the extracted values into a structured table and save it (CSV supports Chinese characters with UTF-8 encoding):

# Create a tidy table
result_table <- tibble(
  PD_Date = pd_date,
  TD_Content = td_text
)

# Save to CSV (adjust the filename as needed)
write.csv(result_table, "extracted_pd_td.csv", row.names = FALSE, fileEncoding = "UTF-8")

Handling Multiple Entries (If Needed)

If your file has multiple separate entries (each with its own PD/TD), split entries by blank lines and loop through them:

# Load purrr for mapping (install first if needed: install.packages("purrr"))
library(purrr)

# Split file into individual entries (separated by blank lines)
entries <- readLines("your_file.txt", warn = FALSE) %>%
  split(cumsum(. == "")) %>%
  lapply(function(x) paste(x[x != ""], collapse = " ")) %>%
  unlist()

# Function to extract PD/TD from a single entry
extract_entry <- function(entry) {
  pd <- str_extract(entry, "PD\\s+(.*?)\\s+(?=[A-Z]{2}\\s)") %>% str_remove("^PD\\s+")
  td <- str_extract(entry, "TD\\s+(.*)") %>% str_remove("^TD\\s+")
  tibble(PD_Date = pd, TD_Content = td)
}

# Apply to all entries and combine into one table
result_table <- map_dfr(entries, extract_entry)

# Save the multi-entry table
write.csv(result_table, "multiple_entries_pd_td.csv", row.names = FALSE, fileEncoding = "UTF-8")

This should cover your use case. If you hit edge cases (like missing tags or unusual spacing), feel free to tweak the regex or reach out for more help!

内容的提问来源于stack exchange,提问作者Beginner

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.22 07:41:30