爬取地方议会网站PDF遇TypeError:字符串索引必须为整数,求解决
Hey Rachel, let's break down what's going wrong with your web scraping code and figure out how to successfully grab those 2018 April-July meeting minutes PDFs.
First, let's look at the errors in your code
Your TypeError: string indices must be integers comes from several key mistakes in how you're using BeautifulSoup and handling links:
You forgot to initialize the main soup object
You defined the base URL but never sent a request to fetch the page content or create the BeautifulSoup instance. Before usingsoup.find(), you need to add:r = requests.get(url) soup = bs(r.text, 'lxml')meeting_linksis a single tag, not a listsoup.find()returns one single HTML tag, not an iterable list. When you try to loop overmeeting_links, you're actually looping through each character in the tag's string content—solinkbecomes a string, and trying to accesslink['text']throws that TypeError because strings only take integer indices. You should usesoup.find_all()to get all relevant meeting links, or first locate the parent container that holds all meeting entries.Date filtering syntax is totally off
- You can't use
link['text']to get a tag's text—uselink.textorlink.get_text()instead. find_all()is a BeautifulSoup method for tags, not strings. To check if a tag's text contains any of your target dates, use a simple check withany():any(date in link.get_text(strip=True) for date in ["April 2018", "May 2018", "June 2018", "July 2018"])
- You can't use
File handling could be safer
Instead of manually opening and closing the file, use awithstatement—it automatically closes the file even if something goes wrong:with open(r"E:\Internship\WORK\GMCA\Getting PDFS\gmcabusinessminutelinks.txt", "w+", encoding='utf-8') as f: # Your write logic here
Yes, you absolutely can filter PDFs by date text!
Here's a revised, working version of your code that implements date filtering correctly. I've also added fixes for relative links and page structure handling:
import requests import time from bs4 import BeautifulSoup as bs # Target committee page URL (skip the base URL and go straight here) committee_url = "https://www.gmcameetings.co.uk/meetings/committee/36/economy_business_growth_and_skills_overview_and_scrutiny" # List of dates we want to target target_dates = ["April 2018", "May 2018", "June 2018", "July 2018"] # Fetch the main committee meeting list page response = requests.get(committee_url) soup = bs(response.text, 'lxml') # Use a with statement for safe file handling with open(r"E:\Internship\WORK\GMCA\Getting PDFS\gmcabusinessminutelinks.txt", "w+", encoding='utf-8') as f: # Find all meeting entries (you may need to adjust this selector using your browser's dev tools) # I'm assuming meetings are wrapped in divs with a class like 'item'—check the actual page HTML! meeting_entries = soup.find_all('div', class_='item') for entry in meeting_entries: # Get the link to the individual meeting page meeting_link = entry.find('a') if not meeting_link: continue # Clean up the meeting text to check for dates meeting_title = meeting_link.get_text(strip=True) # Check if this meeting is in our target date range if any(date in meeting_title for date in target_dates): meeting_href = meeting_link['href'] # Fix relative links by adding the base URL if needed if not meeting_href.startswith('http'): meeting_href = f"https://www.gmcameetings.co.uk{meeting_href}" print(f"Processing meeting: {meeting_title}") # Add a small delay to avoid overwhelming the server time.sleep(1) # Fetch the individual meeting page meeting_response = requests.get(meeting_href) meeting_soup = bs(meeting_response.text, 'lxml') # Find all links that point to minutes PDFs for pdf_link in meeting_soup.find_all('a', href=True): pdf_href = pdf_link['href'] # Check if the link is for minutes and is a PDF if 'minutes' in pdf_href.lower() and pdf_href.endswith('.pdf'): # Fix relative links again if not pdf_href.startswith('http'): pdf_href = f"https://www.gmcameetings.co.uk{pdf_href}" print(f"Found minutes PDF: {pdf_href}") f.write(f"{pdf_href}\n")
A few important notes:
- Adjust the selectors! The
'div', class_='item'selector is a guess—you need to use your browser's developer tools (F12) to inspect the actual HTML structure of the meeting list and update this to match the real tags/classes on the site. - Respect the server! Adding
time.sleep(1)between requests helps avoid getting blocked for too many rapid requests. - Check for relative links: Many site links are relative (don't start with http), so we need to append the base URL to make them usable.
内容的提问来源于stack exchange,提问作者Rachel9866

