使用PyPDF提取PDF文本存入CSV时,如何避免内容拆分至多行
解决PyPDF提取PDF文本存入CSV时多行拆分问题
问题描述
提取长PDF文本存入CSV时,单元格内的文本会被拆分成多行,短PDF则正常。示例输出:
column1 column2 long pdf hello my name is jhone short pdf hello my name is jhone. I haven't any problem for short pdf file
原始代码:
pdf_url ='https://www.snb.ch/en/mmr/speeches/id/ref_20230330_amrtmo/source/ref_20230330_amrtmo.en.pdf' print("pdf_url: ",pdf_url) # Download the PDF file from the URL response = requests.get(pdf_url) # Create an in-memory buffer from the PDF content pdf_buffer = io.BytesIO(response.content) # Read the PDF file from the in-memory buffer pdf = PdfReader(pdf_buffer) pdf_content = [] # Access the contents of the PDF file for page_num in range(len(pdf.pages)): page = pdf.pages[page_num] page = str(page.extract_text()) pdf_content.append(page) with open(filename, "a", newline="", encoding='utf8') as f: writer = csv.writer(f) writer.writerow([first_author, new_date_str, speech_title,pdf_url,pdf_content]) pdf_content.clear()
解决方案
问题根源是PyPDF提取的文本保留了PDF原生的换行符,CSV写入时这些换行符会导致单元格内容换行;同时原始代码将每页文本存入列表,写入CSV时易引发格式异常。需做以下修改:
- 合并所有页面的文本为单个字符串
- 替换所有换行符(
\n、\r)为空格,消除单元格内换行 - 合并连续空格,优化文本紧凑度
修改后的代码:
pdf_url ='https://www.snb.ch/en/mmr/speeches/id/ref_20230330_amrtmo/source/ref_20230330_amrtmo.en.pdf' print("pdf_url: ",pdf_url) # Download the PDF file from the URL response = requests.get(pdf_url) # Create an in-memory buffer from the PDF content pdf_buffer = io.BytesIO(response.content) # Read the PDF file from the in-memory buffer pdf = PdfReader(pdf_buffer) pdf_content = "" # 改为字符串存储所有页面内容 # Access the contents of the PDF file for page_num in range(len(pdf.pages)): page = pdf.pages[page_num] page_text = str(page.extract_text()) # 替换换行符为空格,合并连续空白 cleaned_text = ' '.join(page_text.replace('\r', ' ').split()) pdf_content += cleaned_text + " " # 拼接所有页面文本 with open(filename, "a", newline="", encoding='utf8') as f: writer = csv.writer(f) # 写入处理后的单个字符串,避免格式异常 writer.writerow([first_author, new_date_str, speech_title,pdf_url,pdf_content]) pdf_content = "" # 清空字符串
说明
replace('\r', ' '):处理Windows风格的换行符split()+' '.join():将任意数量的连续空白(包括换行、多空格)替换为单个空格,保证文本无换行且格式整洁- 将
pdf_content改为字符串,避免列表写入CSV时的拆分问题
内容的提问来源于stack exchange,提问作者boyenec
相关产品推荐
相关产品推荐

