使用requests.session爬取网站时遇ValueError: No JSON对象无法解码问题求助
Hey there! Let's break down what's going wrong here and walk through how to fix your web crawler issue.
核心问题:你请求的内容根本不是JSON格式
When you run session.get('http://www.lovetvshow.com/', headers=header), the server sends back the website's full HTML homepage—not a JSON object. The r.json() method in Requests only works when the response content is valid JSON. Since you’re trying to parse HTML as JSON, it’s no wonder you get the "No JSON object could be decoded" error.
You can confirm this easily: replace print r.json() with print r.text, and you’ll see a mess of HTML tags instead of a structured JSON string.
你的代码里的小偏差
Looking at your code, you commented out a POST request for Renren’s login, but then switched to a GET request for lovetvshow.com’s homepage. If your goal is to log into lovetvshow.com, this approach is off-target:
- You need to send your login credentials (
emailandpassword) via a POST request to the site’s actual login API endpoint, not just fetch the homepage.
解决步骤建议
Find the correct login endpoint
- Open your browser’s Developer Tools (F12), go to the "Network" tab, then log into
lovetvshow.commanually. Look for the POST request that sends your login info—this is the endpoint you need to target withsession.post(). - The endpoint might look like
http://www.lovetvshow.com/loginor a similar path, not the root homepage.
- Open your browser’s Developer Tools (F12), go to the "Network" tab, then log into
Check response type before parsing
- Always verify if a response is JSON before calling
r.json(). You can do this by checking theContent-Typeheader:print(r.headers.get('Content-Type')) - If the output is
application/json, thenr.json()is safe to use. If it’stext/html(like your current homepage response), use an HTML parser like BeautifulSoup instead.
- Always verify if a response is JSON before calling
Adjust your code for proper login flow
- Once you have the right login endpoint, update your code to send the POST request there:
session = requests.session() login_data = {'email': email, 'password': password} header = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/63.0.3239.132 Safari/537.36', 'Host': 'www.lovetvshow.com' } # Replace with the actual login endpoint you found login_response = session.post('http://www.lovetvshow.com/your-login-endpoint', data=login_data, headers=header) # Check login success (e.g., status code 200, or look for success flags in the response) print(login_response.status_code) # Parse JSON only if the response is JSON-formatted if login_response.headers.get('Content-Type') == 'application/json': print(login_response.json()) # Else, handle the HTML response (e.g., check if it redirects to a logged-in page) else: print(login_response.text) # After successful login, use the session to access protected pages (cookies are auto-carried) protected_page = session.get('http://www.lovetvshow.com/some-protected-page') print(protected_page.text)
- Once you have the right login endpoint, update your code to send the POST request there:
Handle captchas (your original goal)
- Once you have the basic login flow working, you’ll need to tackle captchas. Common approaches include:
- Using OCR tools (like Tesseract) to solve simple text captchas.
- Integrating third-party services for complex captchas (e.g., reCAPTCHA).
- Extracting the captcha image from the login page, saving it, letting a user solve it manually, then adding the captcha code to your
login_data.
- Once you have the basic login flow working, you’ll need to tackle captchas. Common approaches include:
内容的提问来源于stack exchange,提问作者user3917757

