You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python爬虫批量抓取PDF文本仅返回首个结果故障排查

问题描述

现有代码用于遍历目标主URL(https://covapp.vancouver.ca/councilMeetingPublic/CouncilMeetings.aspx?SearchType=3),点击进入「Council」超链接,通过PyPDF2提取各页面中会议纪要PDF文件的文本内容。
当前异常:代码逻辑设定为循环遍历n个页面批量抓取所有符合要求的PDF,但最终输出仅包含第一个PDF的内容。经校验minutes_links变量已经存储了正确数量的PDF链接,但在执行提取pdf_name、pages_text的for循环时,仅首个链接被处理并存储。

完整代码
import os
import time
from io import BytesIO
from urllib.parse import urljoin

import pandas as pd
import PyPDF2
import requests
from bs4 import BeautifulSoup as soup
from selenium import webdriver
from selenium.webdriver.chrome.options import Options

# Create a headless chromedriver to query and perform action on webpages like a browser
chrome_options = Options()
chrome_options.add_argument("--headless")
driver = webdriver.Chrome(options=chrome_options)

# Main url
my_url = (
    "https://covapp.vancouver.ca/councilMeetingPublic/CouncilMeetings.aspx?SearchType=3"
)


def get_n_first_pages(n: int):
    """Get the html text for the first n pages

    Args:
        n (int): The number of pages we want

    Returns:
        List[str]: A list of html text
    """

    # Initialize the variables containing the pages
    pages = []

    # We query the web page with our chrome driver.
    # This way we can iteratively click on the next link to get all the pages we want
    driver.get(my_url)
    # We append the page source code
    pages.append(driver.page_source)

    # Then for all subsequent pages, we click on next and wait to get the page
    for _ in range(1, n):
        driver.find_element_by_css_selector(
            "#LiverpoolTheme_wt93_block_wtMainContent_RichWidgets_wt132_block_wt28"
        ).click()
        # Wait for the page to load
        time.sleep(1)
        # Append the page
        pages.append(driver.page_source)
    return pages


def get_pdf(link: str):
    """Get the pdf text, per PDF pages, for a given link.

    Args:
        link (str): The link where we can retrieve the PDF

    Returns:
        List[str]: A list containing a string per PDF pages
    """

    # We extract the file name
    pdf_name = link.split("/")[-1].split(".")[0]

    # We get the page containing the PDF link
    # Here we don't need the chrome driver since we don't have to click on the link
    # We can just get the PDF using requests after finding the href
    pdf_link_page = requests.get(link)
    page_soup = soup(pdf_link_page.text, "lxml")
    # We get all <a> tag that have href attribute, then we select only the href
    # containing min.pdf, since we only want the PDF for the minutes
    pdf_link = [
        urljoin(link, l.attrs["href"])
        for l in page_soup.find_all("a", {"href": True})
        if "min.pdf" in l.attrs["href"]
    ]
    # There is only one PDF for the minutes so we get the only element in the list
    pdf_link = pdf_link[0]

    # We get the PDF with requests and then get the PDF bytes
    pdf_bytes = requests.get(pdf_link).content
    # We load the bytes into an in memory file (to avoid saving the PDF on disk)
    p = BytesIO(pdf_bytes)
    p.seek(0, os.SEEK_END)

    # Now we can load our PDF in PyPDF2 from memory
    read_pdf = PyPDF2.PdfFileReader(p)
    count = read_pdf.numPages
    pages_txt = []
    # For each page we extract the text
    for i in range(count):
        page = read_pdf.getPage(i)
        pages_txt.append(page.extractText())

    # We return the PDF name as well as the text inside each pages
    return pdf_name, pages_txt


# Get the first 16 pages, you can change this number
pages = get_n_first_pages(16)


# Initialize a list to store each dataframe rows
df_rows = []

# We iterate over each page
for page in pages:
    page_soup = soup(page, "lxml")

    # Here we get only the <a> tag inside the tbody and each tr
    # We avoid getting the links from the head of the table
    all_links = page_soup.select("tbody tr a")
    # We extract the href for only the links containing council (we don't care about the
    # video link)
    minutes_links = [x.attrs["href"] for x in all_links if "council" in x.attrs["href"]]

    #
    for link in minutes_links:
        pdf_name, pages_text = get_pdf(link)

        df_rows.append(
            {
                "PDF_file_name": pdf_name,
                # We join each page in the list into one string, separting them with a line return
                "PDF_text": "\n".join(pages_text),
            }
        )

        break
    break

# We create the data frame from the list of rows
df = pd.DataFrame(df_rows)
期望输出

最终需要得到包含所有PDF文件名和对应文本的pandas DataFrame,结构示例如下:

PDF_file_namePDF_text
spec20210729min[[' \n \n \n \n \n \n \nSPECIAL COUNCIL MEET\nING MINUTES\n \n \nJULY 29, 2021\n \n \nA Special Meeting of the Council\n \nof the City of Vancouver\n \nw
spec20210802min[[' \n \n \n \n \n \n \nSPECIAL COUNCIL MEET\nING MINUTES\n \n \nAUGUST 2, 2021\n \n \nA Special Meeting of the Council\n \nof the City of Vancouver\n \nw
问题成因
  1. 两处多余的break语句打断了循环执行:
    • 内层遍历minutes_links的循环末尾加了break,每一页仅处理第一个PDF链接就停止遍历当前页剩余的链接
    • 外层遍历pages的循环末尾加了break,仅处理第一页内容就停止遍历所有后续页面
  2. 潜在PDF读取异常:代码中对存储PDF内容的BytesIO对象执行了p.seek(0, os.SEEK_END),将文件指针移到了文件末尾,PyPDF2读取时无法获取到有效内容。
修复方案
  1. 删除代码中两个多余的break语句
  2. 将p.seek(0, os.SEEK_END)修改为p.seek(0),把文件指针重置到PDF内容的开头,保证PyPDF2可以正常读取

修改后的核心逻辑示例:

# 遍历所有页面
for page in pages:
    page_soup = soup(page, "lxml")
    all_links = page_soup.select("tbody tr a")
    minutes_links = [x.attrs["href"] for x in all_links if "council" in x.attrs["href"]]
    # 遍历当前页所有符合要求的PDF链接
    for link in minutes_links:
        pdf_name, pages_text = get_pdf(link)
        df_rows.append(
            {
                "PDF_file_name": pdf_name,
                "PDF_text": "\n".join(pages_text),
            }
        )

# get_pdf函数内的seek逻辑修改
p = BytesIO(pdf_bytes)
p.seek(0)

内容的提问来源于stack exchange,提问作者scotiaboy

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.06 10:54:02