如何在R中按段落拆分PDF内容并导出为CSV?
Great question! Your current code splits on every single line break, which is why it's not respecting natural paragraph boundaries—those are usually marked by one or more blank lines in the PDF's text output. Let's fix that with a simple adjustment to your string splitting logic:
Solution: Split PDF Content by Paragraphs
We'll target blank line separators (instead of individual line breaks) to extract full paragraphs, then carry on with your existing column-splitting workflow. Here's the revised code using the stringr package you're already using:
library(stringr) # Step 1: Extract raw text from the PDF tx <- pdf_text("Name.pdf") # Step 2: Split into paragraphs using blank lines as separators # Regex breakdown: matches one or more blank lines (including \r\n variants with whitespace) paragraphs <- str_split(tx, "\\s*\\n\\s*\\n\\s*")[[1]] # Step 3: Clean up paragraphs (trim whitespace and remove empty entries) clean_paragraphs <- str_trim(paragraphs) clean_paragraphs <- clean_paragraphs[clean_paragraphs != ""] # Step 4: Split each paragraph into 5 columns (matching your original logic) tx3 <- str_split_fixed(clean_paragraphs, "\\s{2,}", 5) # Step 5: Write to CSV write.csv(tx3, file="Path\\ds1.csv", row.names=FALSE)
Key Details
- The regex
\\s*\\n\\s*\\n\\s*is the core fix: it identifies two consecutive newlines (with optional surrounding whitespace) as the paragraph divider, which aligns with how most structured PDFs are formatted. - Using
[[1]]flattens the list output fromstr_split(sincepdf_textreturns one string per page) into a single vector of paragraphs. - The cleanup step removes any stray empty strings that might come from leading/trailing blank pages or extra line breaks in the PDF.
Edge Case Adjustments
If your PDF uses non-standard paragraph separators, tweak the splitting regex:
- Indentation instead of blank lines: Split on newlines followed by 2+ spaces or a tab:
paragraphs <- str_split(tx, "\\n\\s{2,}|\\n\\t")[[1]] - Complex layout: For PDFs where paragraphs span lines without blank separators, you might need to use
pdftools::pdf_data()to extract positional layout data (this requires more advanced parsing to group lines into paragraphs based on their x/y coordinates).
内容的提问来源于stack exchange,提问作者Parul Batra
相关产品推荐
相关产品推荐

