You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用PyPDF提取PDF文本存入CSV时,如何避免内容拆分至多行

解决PyPDF提取PDF文本存入CSV时多行拆分问题

问题描述

提取长PDF文本存入CSV时,单元格内的文本会被拆分成多行,短PDF则正常。示例输出:

column1    column2            
 long pdf    hello my
             name is jhone
 short pdf   hello my name is jhone. I haven't any problem for short pdf file 

原始代码:

pdf_url ='https://www.snb.ch/en/mmr/speeches/id/ref_20230330_amrtmo/source/ref_20230330_amrtmo.en.pdf'
print("pdf_url: ",pdf_url)
   
# Download the PDF file from the URL
response = requests.get(pdf_url)

# Create an in-memory buffer from the PDF content
pdf_buffer = io.BytesIO(response.content)

# Read the PDF file from the in-memory buffer
pdf = PdfReader(pdf_buffer)
pdf_content = []
# Access the contents of the PDF file
for page_num in range(len(pdf.pages)):
    page = pdf.pages[page_num]
    page = str(page.extract_text())
    pdf_content.append(page)
    
   
with open(filename, "a", newline="",  encoding='utf8') as f:
        writer = csv.writer(f)
        writer.writerow([first_author, new_date_str, speech_title,pdf_url,pdf_content])

pdf_content.clear()  

解决方案

问题根源是PyPDF提取的文本保留了PDF原生的换行符,CSV写入时这些换行符会导致单元格内容换行;同时原始代码将每页文本存入列表,写入CSV时易引发格式异常。需做以下修改:

  • 合并所有页面的文本为单个字符串
  • 替换所有换行符(\n、\r)为空格,消除单元格内换行
  • 合并连续空格,优化文本紧凑度

修改后的代码:

pdf_url ='https://www.snb.ch/en/mmr/speeches/id/ref_20230330_amrtmo/source/ref_20230330_amrtmo.en.pdf'
print("pdf_url: ",pdf_url)
   
# Download the PDF file from the URL
response = requests.get(pdf_url)

# Create an in-memory buffer from the PDF content
pdf_buffer = io.BytesIO(response.content)

# Read the PDF file from the in-memory buffer
pdf = PdfReader(pdf_buffer)
pdf_content = ""  # 改为字符串存储所有页面内容
# Access the contents of the PDF file
for page_num in range(len(pdf.pages)):
    page = pdf.pages[page_num]
    page_text = str(page.extract_text())
    # 替换换行符为空格,合并连续空白
    cleaned_text = ' '.join(page_text.replace('\r', ' ').split())
    pdf_content += cleaned_text + " "  # 拼接所有页面文本
    
   
with open(filename, "a", newline="",  encoding='utf8') as f:
        writer = csv.writer(f)
        # 写入处理后的单个字符串,避免格式异常
        writer.writerow([first_author, new_date_str, speech_title,pdf_url,pdf_content])

pdf_content = ""  # 清空字符串

说明

  • replace('\r', ' '):处理Windows风格的换行符
  • split() + ' '.join():将任意数量的连续空白(包括换行、多空格)替换为单个空格,保证文本无换行且格式整洁
  • 将pdf_content改为字符串,避免列表写入CSV时的拆分问题

内容的提问来源于stack exchange,提问作者boyenec

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.22 10:13:29