You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

爬取地方议会网站PDF遇TypeError:字符串索引必须为整数,求解决

解决GMCA议会网站爬取PDF时的TypeError问题及日期筛选方案

Hey Rachel, let's break down what's going wrong with your web scraping code and figure out how to successfully grab those 2018 April-July meeting minutes PDFs.

First, let's look at the errors in your code

Your TypeError: string indices must be integers comes from several key mistakes in how you're using BeautifulSoup and handling links:

  1. You forgot to initialize the main soup object
    You defined the base URL but never sent a request to fetch the page content or create the BeautifulSoup instance. Before using soup.find(), you need to add:

    r = requests.get(url)
    soup = bs(r.text, 'lxml')
    
  2. meeting_links is a single tag, not a list
    soup.find() returns one single HTML tag, not an iterable list. When you try to loop over meeting_links, you're actually looping through each character in the tag's string content—so link becomes a string, and trying to access link['text'] throws that TypeError because strings only take integer indices. You should use soup.find_all() to get all relevant meeting links, or first locate the parent container that holds all meeting entries.

  3. Date filtering syntax is totally off

    • You can't use link['text'] to get a tag's text—use link.text or link.get_text() instead.
    • find_all() is a BeautifulSoup method for tags, not strings. To check if a tag's text contains any of your target dates, use a simple check with any():
      any(date in link.get_text(strip=True) for date in ["April 2018", "May 2018", "June 2018", "July 2018"])
      
  4. File handling could be safer
    Instead of manually opening and closing the file, use a with statement—it automatically closes the file even if something goes wrong:

    with open(r"E:\Internship\WORK\GMCA\Getting PDFS\gmcabusinessminutelinks.txt", "w+", encoding='utf-8') as f:
        # Your write logic here
    

Yes, you absolutely can filter PDFs by date text!

Here's a revised, working version of your code that implements date filtering correctly. I've also added fixes for relative links and page structure handling:

import requests
import time
from bs4 import BeautifulSoup as bs

# Target committee page URL (skip the base URL and go straight here)
committee_url = "https://www.gmcameetings.co.uk/meetings/committee/36/economy_business_growth_and_skills_overview_and_scrutiny"
# List of dates we want to target
target_dates = ["April 2018", "May 2018", "June 2018", "July 2018"]

# Fetch the main committee meeting list page
response = requests.get(committee_url)
soup = bs(response.text, 'lxml')

# Use a with statement for safe file handling
with open(r"E:\Internship\WORK\GMCA\Getting PDFS\gmcabusinessminutelinks.txt", "w+", encoding='utf-8') as f:
    # Find all meeting entries (you may need to adjust this selector using your browser's dev tools)
    # I'm assuming meetings are wrapped in divs with a class like 'item'—check the actual page HTML!
    meeting_entries = soup.find_all('div', class_='item')
    
    for entry in meeting_entries:
        # Get the link to the individual meeting page
        meeting_link = entry.find('a')
        if not meeting_link:
            continue
        
        # Clean up the meeting text to check for dates
        meeting_title = meeting_link.get_text(strip=True)
        
        # Check if this meeting is in our target date range
        if any(date in meeting_title for date in target_dates):
            meeting_href = meeting_link['href']
            # Fix relative links by adding the base URL if needed
            if not meeting_href.startswith('http'):
                meeting_href = f"https://www.gmcameetings.co.uk{meeting_href}"
            
            print(f"Processing meeting: {meeting_title}")
            # Add a small delay to avoid overwhelming the server
            time.sleep(1)
            
            # Fetch the individual meeting page
            meeting_response = requests.get(meeting_href)
            meeting_soup = bs(meeting_response.text, 'lxml')
            
            # Find all links that point to minutes PDFs
            for pdf_link in meeting_soup.find_all('a', href=True):
                pdf_href = pdf_link['href']
                # Check if the link is for minutes and is a PDF
                if 'minutes' in pdf_href.lower() and pdf_href.endswith('.pdf'):
                    # Fix relative links again
                    if not pdf_href.startswith('http'):
                        pdf_href = f"https://www.gmcameetings.co.uk{pdf_href}"
                    print(f"Found minutes PDF: {pdf_href}")
                    f.write(f"{pdf_href}\n")

A few important notes:

  • Adjust the selectors! The 'div', class_='item' selector is a guess—you need to use your browser's developer tools (F12) to inspect the actual HTML structure of the meeting list and update this to match the real tags/classes on the site.
  • Respect the server! Adding time.sleep(1) between requests helps avoid getting blocked for too many rapid requests.
  • Check for relative links: Many site links are relative (don't start with http), so we need to append the base URL to make them usable.

内容的提问来源于stack exchange,提问作者Rachel9866

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.13 09:18:51