Python爬虫批量抓取PDF文本仅返回首个结果故障排查
问题描述
现有代码用于遍历目标主URL(https://covapp.vancouver.ca/councilMeetingPublic/CouncilMeetings.aspx?SearchType=3),点击进入「Council」超链接,通过PyPDF2提取各页面中会议纪要PDF文件的文本内容。
当前异常:代码逻辑设定为循环遍历n个页面批量抓取所有符合要求的PDF,但最终输出仅包含第一个PDF的内容。经校验minutes_links变量已经存储了正确数量的PDF链接,但在执行提取pdf_name、pages_text的for循环时,仅首个链接被处理并存储。
完整代码
import os import time from io import BytesIO from urllib.parse import urljoin import pandas as pd import PyPDF2 import requests from bs4 import BeautifulSoup as soup from selenium import webdriver from selenium.webdriver.chrome.options import Options # Create a headless chromedriver to query and perform action on webpages like a browser chrome_options = Options() chrome_options.add_argument("--headless") driver = webdriver.Chrome(options=chrome_options) # Main url my_url = ( "https://covapp.vancouver.ca/councilMeetingPublic/CouncilMeetings.aspx?SearchType=3" ) def get_n_first_pages(n: int): """Get the html text for the first n pages Args: n (int): The number of pages we want Returns: List[str]: A list of html text """ # Initialize the variables containing the pages pages = [] # We query the web page with our chrome driver. # This way we can iteratively click on the next link to get all the pages we want driver.get(my_url) # We append the page source code pages.append(driver.page_source) # Then for all subsequent pages, we click on next and wait to get the page for _ in range(1, n): driver.find_element_by_css_selector( "#LiverpoolTheme_wt93_block_wtMainContent_RichWidgets_wt132_block_wt28" ).click() # Wait for the page to load time.sleep(1) # Append the page pages.append(driver.page_source) return pages def get_pdf(link: str): """Get the pdf text, per PDF pages, for a given link. Args: link (str): The link where we can retrieve the PDF Returns: List[str]: A list containing a string per PDF pages """ # We extract the file name pdf_name = link.split("/")[-1].split(".")[0] # We get the page containing the PDF link # Here we don't need the chrome driver since we don't have to click on the link # We can just get the PDF using requests after finding the href pdf_link_page = requests.get(link) page_soup = soup(pdf_link_page.text, "lxml") # We get all <a> tag that have href attribute, then we select only the href # containing min.pdf, since we only want the PDF for the minutes pdf_link = [ urljoin(link, l.attrs["href"]) for l in page_soup.find_all("a", {"href": True}) if "min.pdf" in l.attrs["href"] ] # There is only one PDF for the minutes so we get the only element in the list pdf_link = pdf_link[0] # We get the PDF with requests and then get the PDF bytes pdf_bytes = requests.get(pdf_link).content # We load the bytes into an in memory file (to avoid saving the PDF on disk) p = BytesIO(pdf_bytes) p.seek(0, os.SEEK_END) # Now we can load our PDF in PyPDF2 from memory read_pdf = PyPDF2.PdfFileReader(p) count = read_pdf.numPages pages_txt = [] # For each page we extract the text for i in range(count): page = read_pdf.getPage(i) pages_txt.append(page.extractText()) # We return the PDF name as well as the text inside each pages return pdf_name, pages_txt # Get the first 16 pages, you can change this number pages = get_n_first_pages(16) # Initialize a list to store each dataframe rows df_rows = [] # We iterate over each page for page in pages: page_soup = soup(page, "lxml") # Here we get only the <a> tag inside the tbody and each tr # We avoid getting the links from the head of the table all_links = page_soup.select("tbody tr a") # We extract the href for only the links containing council (we don't care about the # video link) minutes_links = [x.attrs["href"] for x in all_links if "council" in x.attrs["href"]] # for link in minutes_links: pdf_name, pages_text = get_pdf(link) df_rows.append( { "PDF_file_name": pdf_name, # We join each page in the list into one string, separting them with a line return "PDF_text": "\n".join(pages_text), } ) break break # We create the data frame from the list of rows df = pd.DataFrame(df_rows)
期望输出
最终需要得到包含所有PDF文件名和对应文本的pandas DataFrame,结构示例如下:
| PDF_file_name | PDF_text |
|---|---|
| spec20210729min | [[' \n \n \n \n \n \n \nSPECIAL COUNCIL MEET\nING MINUTES\n \n \nJULY 29, 2021\n \n \nA Special Meeting of the Council\n \nof the City of Vancouver\n \nw |
| spec20210802min | [[' \n \n \n \n \n \n \nSPECIAL COUNCIL MEET\nING MINUTES\n \n \nAUGUST 2, 2021\n \n \nA Special Meeting of the Council\n \nof the City of Vancouver\n \nw |
问题成因
- 两处多余的
break语句打断了循环执行:- 内层遍历
minutes_links的循环末尾加了break,每一页仅处理第一个PDF链接就停止遍历当前页剩余的链接 - 外层遍历
pages的循环末尾加了break,仅处理第一页内容就停止遍历所有后续页面
- 内层遍历
- 潜在PDF读取异常:代码中对存储PDF内容的BytesIO对象执行了
p.seek(0, os.SEEK_END),将文件指针移到了文件末尾,PyPDF2读取时无法获取到有效内容。
修复方案
- 删除代码中两个多余的
break语句 - 将
p.seek(0, os.SEEK_END)修改为p.seek(0),把文件指针重置到PDF内容的开头,保证PyPDF2可以正常读取
修改后的核心逻辑示例:
# 遍历所有页面 for page in pages: page_soup = soup(page, "lxml") all_links = page_soup.select("tbody tr a") minutes_links = [x.attrs["href"] for x in all_links if "council" in x.attrs["href"]] # 遍历当前页所有符合要求的PDF链接 for link in minutes_links: pdf_name, pages_text = get_pdf(link) df_rows.append( { "PDF_file_name": pdf_name, "PDF_text": "\n".join(pages_text), } ) # get_pdf函数内的seek逻辑修改 p = BytesIO(pdf_bytes) p.seek(0)
内容的提问来源于stack exchange,提问作者scotiaboy
相关产品推荐
相关产品推荐

