You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python 3爬取需登录网站(censos.ccp.ucr.ac.cr)求助

Troubleshooting Your Python Login Scraper for censos.ccp.ucr.ac.cr

Hey there! As a Python scraping newbie, it’s totally normal to hit snags with authenticated pages—let’s walk through the most common issues that might be breaking your script, and how to fix them.

Common Issues & Fixes

1. You’re Not Maintaining a Session

Most login systems use cookies to track authenticated users. If you’re just using requests.post() directly without a session object, the cookies from the login request won’t carry over to subsequent requests, meaning you’ll stay logged out no matter what.

Fix: Use requests.Session()—this object automatically handles cookie persistence for you across all requests:

import requests
session = requests.Session()  # All requests through this session share cookies

2. You’re Missing Hidden Form Fields

Login pages often include hidden fields (like CSRF tokens) that the server requires to validate your login request. You found codigo and contrasena, but there might be other hidden inputs you’re not including in your POST data.

Fix: First fetch the login page to extract all form fields, including hidden ones:

from bs4 import BeautifulSoup

login_url = "https://censos.ccp.ucr.ac.cr/"
# Mimic a real browser's user agent to avoid being blocked
headers = {"User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36"}

# Get the login page to scrape hidden fields
response = session.get(login_url, headers=headers)
soup = BeautifulSoup(response.text, "html.parser")

# Build your login data with username, password, and hidden fields
login_data = {
    "codigo": "your_username_here",
    "contrasena": "your_password_here"
}

# Add all hidden input fields to the data dict
for hidden_input in soup.find_all("input", type="hidden"):
    input_name = hidden_input.get("name")
    input_value = hidden_input.get("value")
    if input_name and input_value:
        login_data[input_name] = input_value

3. Your Request Headers Are Too Generic

Servers often block requests that don’t look like they’re coming from a real browser. Missing headers like User-Agent or Referer can cause your login request to be rejected outright.

Fix: Include realistic headers in your requests—like the User-Agent in the code above—to mimic a browser’s behavior.

4. You’re Posting to the Wrong URL

Double-check the <form> tag’s action attribute on the login page. Sometimes the form submits to a different URL than the page you’re viewing (e.g., /login instead of the root URL). If action has a specific path, your POST request should target that full URL.

Full Example Script

Putting it all together, here’s a complete script to test:

import requests
from bs4 import BeautifulSoup

def login_to_censos(username, password):
    session = requests.Session()
    headers = {
        "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36",
        "Referer": "https://censos.ccp.ucr.ac.cr/"
    }

    # Fetch login page to get hidden fields
    login_page = session.get("https://censos.ccp.ucr.ac.cr/", headers=headers)
    soup = BeautifulSoup(login_page.text, "html.parser")

    # Prepare login data
    login_data = {"codigo": username, "contrasena": password}
    for hidden in soup.find_all("input", type="hidden"):
        if hidden.get("name"):
            login_data[hidden["name"]] = hidden["value"]

    # Submit login request (update URL if form action is different)
    login_response = session.post(
        "https://censos.ccp.ucr.ac.cr/",
        data=login_data,
        headers=headers
    )

    # Verify login by checking a protected page (replace with your target path)
    protected_page = session.get("https://censos.ccp.ucr.ac.cr/protected_path", headers=headers)
    return protected_page.text

# Test with your credentials
print(login_to_censos("your_codigo", "your_contrasena"))

Final Checks

  • If you still get an error, check login_response.status_code (should be 200 or 302 if redirected after login).
  • Look out for CAPTCHAs—if the site uses them, you’ll need manual input or specialized tools to get past them.
  • Inspect the actual login request in your browser’s DevTools (Network tab) to see exactly what data is being sent, and match that in your script.

内容的提问来源于stack exchange,提问作者chapunosky

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 09:35:20