IEEE论文关键词爬虫触发IndexError,求修复方案(附代码)
Fixing IndexError When Scraping IEEE Xplore Keywords
Hey there! Let's figure out why you're hitting that IndexError: list index out of range and get your keyword scraper working again.
What's Causing the Error?
The error happens because re.findall('global.document.metadata=(.*;)', i)[0] tries to access the first element of an empty list. That means your regex isn't finding any matches for global.document.metadata=(.*;) in the 9th script tag. A few common reasons for this:
- IEEE Xplore's page structure might have changed, so the metadata isn't in the 9th script tag anymore.
- Your regex pattern is too rigid—maybe the assignment syntax has shifted (e.g., extra spaces, no trailing semicolon, or different quote usage).
- You forgot to import the
remodule (it's missing in your code snippet, which would cause an error too!).
Step-by-Step Fix
Here's a revised version of your code that addresses these issues:
import requests import json import re # Don't forget this import! from bs4 import BeautifulSoup # Add headers to mimic a real browser (helps avoid anti-scraping blocks) headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36' } ieee_content = requests.get("http://ieeexplore.ieee.org/document/8465981", headers=headers, timeout=180) # Use html.parser instead of xml—better for parsing HTML pages soup = BeautifulSoup(ieee_content.text, 'html.parser') metadata = None # Loop through all script tags instead of hardcoding index 9 for script_tag in soup.find_all('script'): script_text = script_tag.string # Skip empty script tags if not script_text: continue # Check if the script contains the metadata we need if 'global.document.metadata' in script_text: # More flexible regex to match the metadata assignment # re.DOTALL lets . match newlines (metadata is often a multi-line object) match = re.search(r'global.document.metadata\s*=\s*(.*?);', script_text, re.DOTALL) if match: raw_metadata = match.group(1) # Clean up the string to make it valid JSON cleaned_metadata = raw_metadata.replace("'", '"').rstrip(',') # Remove trailing commas (invalid in JSON) try: metadata = json.loads(cleaned_metadata) break # Exit loop once we find valid metadata except json.JSONDecodeError as e: print(f"Failed to parse JSON: {e}") continue if metadata: # Extract keywords—adjust the key based on actual metadata structure keywords = metadata.get('keywords', []) print("Extracted Keywords:", keywords) else: print("Couldn't find metadata. IEEE's page structure may have updated.")
Key Improvements
- No hardcoded script index: We loop through all script tags and look for the one containing
global.document.metadata, so we don't break if the page structure changes. - Flexible regex: Uses
\s*to match extra spaces around the=sign, andre.DOTALLto handle multi-line metadata objects. - JSON cleanup: Handles single quotes (converts to double quotes) and trailing commas (which JSON doesn't allow) to avoid parsing errors.
- Anti-scraping mitigation: Adds a
User-Agentheader to mimic a real browser, reducing the chance of being blocked. - Error handling: Catches JSON parsing errors and skips invalid script tags instead of crashing.
Extra Notes
- IEEE Xplore has anti-scraping measures, so avoid making too many requests in a short time. Consider adding delays between requests.
- If this still doesn't work, check the page source manually to see if the metadata is now stored in a different format (e.g., a JSON-LD tag or a different script variable name).
内容的提问来源于stack exchange,提问作者Shukuang Chen
相关产品推荐
相关产品推荐

