You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

未使用Session的Requests脚本解析Compranet网页内容失败的修复方法咨询

Fixing Your Non-Session Compranet Scraper

Let's break down the issues in your script and fix them step by step:

Key Problems in the Original Code

  1. Unsupported CSS Selector: BeautifulSoup's html.parser doesn't support the :contains() CSS pseudo-class. This was causing the first error when trying to grab the "Vigentes" link.
  2. Global Header Pollution: You're modifying the global headers dictionary in get_content(), which means subsequent requests will reuse the last Cookie value instead of the correct one for each detail page.
  3. Incomplete Cookie Handling: You're only passing the JSESSIONID cookie, but the site likely requires other cookies (like session-related flags) to validate your session properly.
  4. Fragile Selector Dependencies: Relying on :contains() for field extraction makes the scraper brittle if the site's HTML structure changes slightly.

Modified Script with Fixes

import re
import requests
from bs4 import BeautifulSoup
from urllib.parse import urljoin

base = 'https://compranet.hacienda.gob.mx'
link = 'https://compranet.hacienda.gob.mx/web/login.html'
vigen_detail_page = 'https://compranet.hacienda.gob.mx/esop/toolkit/opportunity/current/{}/detail.si'

# Base headers (never modify this globally)
BASE_HEADERS = {
    'User-Agent': 'Mozilla/5.0 (Windows NT 6.1) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/88.0.4324.150 Safari/537.36',
    'X-Requested-With': 'XMLHttpRequest',
}

def grab_first_link_from_dropdown(link):
    r = requests.get(link, headers=BASE_HEADERS)
    soup = BeautifulSoup(r.text, "html.parser")
    # Replace :contains() with manual text matching since html.parser doesn't support it
    vigentes_link = None
    for a_tag in soup.select('ul.dropdown-menu > li > a'):
        if 'Vigentes' in a_tag.get_text(strip=True):
            vigentes_link = a_tag.get("href")
            break
    if not vigentes_link:
        raise ValueError("Could not find 'Vigentes' link in dropdown")
    return urljoin(base, vigentes_link)

def fetch_detail_page_link(cat_link):
    res = requests.get(cat_link, headers=BASE_HEADERS)
    # Convert all cookies to a proper Cookie string instead of just JSESSIONID
    cookie_str = '; '.join([f'{name}={value}' for name, value in res.cookies.items()])
    soup = BeautifulSoup(res.text, "html.parser")
    for items in soup.select("table.list-table > tbody.list-tbody > tr"):
        target_link = items.select_one("a.detailLink").get("onclick")
        detail_num = re.findall(r"goToDetail\('(\d+?)'", target_link)[0]
        inner_link = vigen_detail_page.format(detail_num)
        yield cookie_str, inner_link

def get_content(cookie_str, inner_link):
    # Create a copy of base headers and add the cookie for this request only
    request_headers = BASE_HEADERS.copy()
    request_headers['Cookie'] = cookie_str
    res = requests.get(inner_link, headers=request_headers)
    soup = BeautifulSoup(res.text, "html.parser")
    
    # More robust way to find the fields without relying on :contains()
    expediente = ""
    descripcion = ""
    for question, answer in zip(soup.select(".form_question"), soup.select(".form_answer")):
        q_text = question.get_text(strip=True)
        if q_text == 'Código del Expediente':
            expediente = answer.get_text(strip=True)
        elif q_text == 'Descripción del Expediente':
            descripcion = answer.get_text(strip=True)
    return expediente, descripcion

if __name__ == '__main__':
    try:
        category_link = grab_first_link_from_dropdown(link)
        for cookie, detail_page_link in fetch_detail_page_link(category_link):
            print(get_content(cookie, detail_page_link))
    except Exception as e:
        print(f"Error occurred: {str(e)}")

What Changed?

  • Replaced :contains(): We now loop through dropdown links and check text directly, since html.parser doesn't support the CSS pseudo-class.
  • Isolated Headers: Instead of modifying global headers, we create a copy for each detail page request to prevent cookie leakage between requests.
  • Full Cookie String: We pass all cookies from the category page response instead of just JSESSIONID—this ensures the server recognizes your session properly.
  • Robust Field Extraction: We pair each question with its answer and match exact text, making the scraper less fragile to minor HTML changes.
  • Basic Error Handling: Added a try/except block in the main function to catch and report execution errors.

Why This Works Without requests.Session

By manually capturing and passing all cookies from the initial category page request to each detail page request, we replicate the session persistence that requests.Session handles automatically. The key was ensuring we didn't miss any required cookies and didn't pollute the global header state.

内容的提问来源于stack exchange,提问作者robots.txt

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.30 18:27:43