You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在R中按段落拆分PDF内容并导出为CSV?

Great question! Your current code splits on every single line break, which is why it's not respecting natural paragraph boundaries—those are usually marked by one or more blank lines in the PDF's text output. Let's fix that with a simple adjustment to your string splitting logic:

Solution: Split PDF Content by Paragraphs

We'll target blank line separators (instead of individual line breaks) to extract full paragraphs, then carry on with your existing column-splitting workflow. Here's the revised code using the stringr package you're already using:

library(stringr)

# Step 1: Extract raw text from the PDF
tx <- pdf_text("Name.pdf")

# Step 2: Split into paragraphs using blank lines as separators
# Regex breakdown: matches one or more blank lines (including \r\n variants with whitespace)
paragraphs <- str_split(tx, "\\s*\\n\\s*\\n\\s*")[[1]]

# Step 3: Clean up paragraphs (trim whitespace and remove empty entries)
clean_paragraphs <- str_trim(paragraphs)
clean_paragraphs <- clean_paragraphs[clean_paragraphs != ""]

# Step 4: Split each paragraph into 5 columns (matching your original logic)
tx3 <- str_split_fixed(clean_paragraphs, "\\s{2,}", 5)

# Step 5: Write to CSV
write.csv(tx3, file="Path\\ds1.csv", row.names=FALSE)

Key Details

  • The regex \\s*\\n\\s*\\n\\s* is the core fix: it identifies two consecutive newlines (with optional surrounding whitespace) as the paragraph divider, which aligns with how most structured PDFs are formatted.
  • Using [[1]] flattens the list output from str_split (since pdf_text returns one string per page) into a single vector of paragraphs.
  • The cleanup step removes any stray empty strings that might come from leading/trailing blank pages or extra line breaks in the PDF.

Edge Case Adjustments

If your PDF uses non-standard paragraph separators, tweak the splitting regex:

  • Indentation instead of blank lines: Split on newlines followed by 2+ spaces or a tab:
    paragraphs <- str_split(tx, "\\n\\s{2,}|\\n\\t")[[1]]
    
  • Complex layout: For PDFs where paragraphs span lines without blank separators, you might need to use pdftools::pdf_data() to extract positional layout data (this requires more advanced parsing to group lines into paragraphs based on their x/y coordinates).

内容的提问来源于stack exchange,提问作者Parul Batra

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.14 08:00:37