You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

IEEE论文关键词爬虫触发IndexError,求修复方案(附代码)

Fixing IndexError When Scraping IEEE Xplore Keywords

Hey there! Let's figure out why you're hitting that IndexError: list index out of range and get your keyword scraper working again.

What's Causing the Error?

The error happens because re.findall('global.document.metadata=(.*;)', i)[0] tries to access the first element of an empty list. That means your regex isn't finding any matches for global.document.metadata=(.*;) in the 9th script tag. A few common reasons for this:

  • IEEE Xplore's page structure might have changed, so the metadata isn't in the 9th script tag anymore.
  • Your regex pattern is too rigid—maybe the assignment syntax has shifted (e.g., extra spaces, no trailing semicolon, or different quote usage).
  • You forgot to import the re module (it's missing in your code snippet, which would cause an error too!).

Step-by-Step Fix

Here's a revised version of your code that addresses these issues:

import requests
import json
import re  # Don't forget this import!
from bs4 import BeautifulSoup

# Add headers to mimic a real browser (helps avoid anti-scraping blocks)
headers = {
    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36'
}

ieee_content = requests.get("http://ieeexplore.ieee.org/document/8465981", headers=headers, timeout=180)
# Use html.parser instead of xml—better for parsing HTML pages
soup = BeautifulSoup(ieee_content.text, 'html.parser')

metadata = None
# Loop through all script tags instead of hardcoding index 9
for script_tag in soup.find_all('script'):
    script_text = script_tag.string
    # Skip empty script tags
    if not script_text:
        continue
    # Check if the script contains the metadata we need
    if 'global.document.metadata' in script_text:
        # More flexible regex to match the metadata assignment
        # re.DOTALL lets . match newlines (metadata is often a multi-line object)
        match = re.search(r'global.document.metadata\s*=\s*(.*?);', script_text, re.DOTALL)
        if match:
            raw_metadata = match.group(1)
            # Clean up the string to make it valid JSON
            cleaned_metadata = raw_metadata.replace("'", '"').rstrip(',')  # Remove trailing commas (invalid in JSON)
            try:
                metadata = json.loads(cleaned_metadata)
                break  # Exit loop once we find valid metadata
            except json.JSONDecodeError as e:
                print(f"Failed to parse JSON: {e}")
                continue

if metadata:
    # Extract keywords—adjust the key based on actual metadata structure
    keywords = metadata.get('keywords', [])
    print("Extracted Keywords:", keywords)
else:
    print("Couldn't find metadata. IEEE's page structure may have updated.")

Key Improvements

  • No hardcoded script index: We loop through all script tags and look for the one containing global.document.metadata, so we don't break if the page structure changes.
  • Flexible regex: Uses \s* to match extra spaces around the = sign, and re.DOTALL to handle multi-line metadata objects.
  • JSON cleanup: Handles single quotes (converts to double quotes) and trailing commas (which JSON doesn't allow) to avoid parsing errors.
  • Anti-scraping mitigation: Adds a User-Agent header to mimic a real browser, reducing the chance of being blocked.
  • Error handling: Catches JSON parsing errors and skips invalid script tags instead of crashing.

Extra Notes

  • IEEE Xplore has anti-scraping measures, so avoid making too many requests in a short time. Consider adding delays between requests.
  • If this still doesn't work, check the page source manually to see if the metadata is now stored in a different format (e.g., a JSON-LD tag or a different script variable name).

内容的提问来源于stack exchange,提问作者Shukuang Chen

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.12 03:43:45