未使用Session的Requests脚本解析Compranet网页内容失败的修复方法咨询
Fixing Your Non-Session Compranet Scraper
Let's break down the issues in your script and fix them step by step:
Key Problems in the Original Code
- Unsupported CSS Selector: BeautifulSoup's
html.parserdoesn't support the:contains()CSS pseudo-class. This was causing the first error when trying to grab the "Vigentes" link. - Global Header Pollution: You're modifying the global
headersdictionary inget_content(), which means subsequent requests will reuse the last Cookie value instead of the correct one for each detail page. - Incomplete Cookie Handling: You're only passing the
JSESSIONIDcookie, but the site likely requires other cookies (like session-related flags) to validate your session properly. - Fragile Selector Dependencies: Relying on
:contains()for field extraction makes the scraper brittle if the site's HTML structure changes slightly.
Modified Script with Fixes
import re import requests from bs4 import BeautifulSoup from urllib.parse import urljoin base = 'https://compranet.hacienda.gob.mx' link = 'https://compranet.hacienda.gob.mx/web/login.html' vigen_detail_page = 'https://compranet.hacienda.gob.mx/esop/toolkit/opportunity/current/{}/detail.si' # Base headers (never modify this globally) BASE_HEADERS = { 'User-Agent': 'Mozilla/5.0 (Windows NT 6.1) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/88.0.4324.150 Safari/537.36', 'X-Requested-With': 'XMLHttpRequest', } def grab_first_link_from_dropdown(link): r = requests.get(link, headers=BASE_HEADERS) soup = BeautifulSoup(r.text, "html.parser") # Replace :contains() with manual text matching since html.parser doesn't support it vigentes_link = None for a_tag in soup.select('ul.dropdown-menu > li > a'): if 'Vigentes' in a_tag.get_text(strip=True): vigentes_link = a_tag.get("href") break if not vigentes_link: raise ValueError("Could not find 'Vigentes' link in dropdown") return urljoin(base, vigentes_link) def fetch_detail_page_link(cat_link): res = requests.get(cat_link, headers=BASE_HEADERS) # Convert all cookies to a proper Cookie string instead of just JSESSIONID cookie_str = '; '.join([f'{name}={value}' for name, value in res.cookies.items()]) soup = BeautifulSoup(res.text, "html.parser") for items in soup.select("table.list-table > tbody.list-tbody > tr"): target_link = items.select_one("a.detailLink").get("onclick") detail_num = re.findall(r"goToDetail\('(\d+?)'", target_link)[0] inner_link = vigen_detail_page.format(detail_num) yield cookie_str, inner_link def get_content(cookie_str, inner_link): # Create a copy of base headers and add the cookie for this request only request_headers = BASE_HEADERS.copy() request_headers['Cookie'] = cookie_str res = requests.get(inner_link, headers=request_headers) soup = BeautifulSoup(res.text, "html.parser") # More robust way to find the fields without relying on :contains() expediente = "" descripcion = "" for question, answer in zip(soup.select(".form_question"), soup.select(".form_answer")): q_text = question.get_text(strip=True) if q_text == 'Código del Expediente': expediente = answer.get_text(strip=True) elif q_text == 'Descripción del Expediente': descripcion = answer.get_text(strip=True) return expediente, descripcion if __name__ == '__main__': try: category_link = grab_first_link_from_dropdown(link) for cookie, detail_page_link in fetch_detail_page_link(category_link): print(get_content(cookie, detail_page_link)) except Exception as e: print(f"Error occurred: {str(e)}")
What Changed?
- Replaced
:contains(): We now loop through dropdown links and check text directly, sincehtml.parserdoesn't support the CSS pseudo-class. - Isolated Headers: Instead of modifying global headers, we create a copy for each detail page request to prevent cookie leakage between requests.
- Full Cookie String: We pass all cookies from the category page response instead of just
JSESSIONID—this ensures the server recognizes your session properly. - Robust Field Extraction: We pair each question with its answer and match exact text, making the scraper less fragile to minor HTML changes.
- Basic Error Handling: Added a try/except block in the main function to catch and report execution errors.
Why This Works Without requests.Session
By manually capturing and passing all cookies from the initial category page request to each detail page request, we replicate the session persistence that requests.Session handles automatically. The key was ensuring we didn't miss any required cookies and didn't pollute the global header state.
内容的提问来源于stack exchange,提问作者robots.txt
相关产品推荐
相关产品推荐

