You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

用pandas爬HTML表格报错,requests/BeautifulSoup无返回,求排查原因

Common Causes & Fixes for "No text parsed from document" Error

Let’s break down the most likely reasons your scraping code is failing on this specific site—especially since it works elsewhere:

1. You’re not maintaining your logged-in session

Your code sends a login POST request, but then makes a separate request for the table without carrying over the login cookies. Requests uses a fresh session for each call by default, so the server doesn’t recognize you as authenticated—and might return an empty response or redirect you back to the login page.

Fix: Use requests.Session() to preserve your login state across requests:

session = requests.Session()
urllogin = 'http://website.html'
values = {'user': 'id', 'password': 'pass'}
# Log in with the persistent session
session.post(urllogin, data=values)
# Request the table page using the same logged-in session
response = session.get('http://website/table')
# Verify content exists before parsing
print(response.text[:500])  # Print first 500 characters to check
tables = pd.read_html(response.text)

2. Anti-scraping measures are blocking you

Many sites block non-browser requests by checking headers like User-Agent. If your request doesn’t have a valid browser-like user agent, the server might return an empty page instead of the actual content. Other common checks include Referer headers or cookie validation.

Fix: Add a realistic user agent to your requests:

headers = {
    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36'
}
# Include headers in your session request
response = session.get('http://website/table', headers=headers)

3. The table is dynamically loaded with JavaScript

If the table is rendered client-side (e.g., fetched via an API after the initial page loads), tools like requests and BeautifulSoup can’t access it—they only retrieve the static initial HTML, which might be empty or lack the table content entirely.

Fix: Use a tool that renders JavaScript, like Selenium or Playwright. Here’s a quick Selenium example:

from selenium import webdriver
from selenium.webdriver.common.by import By
import pandas as pd

driver = webdriver.Chrome()
# Log in manually via the browser
driver.get('http://website.html')
driver.find_element(By.NAME, 'user').send_keys('id')
driver.find_element(By.NAME, 'password').send_keys('pass')
driver.find_element(By.CSS_SELECTOR, 'button[type="submit"]').click()
# Navigate to the table page
driver.get('http://website/table')
# Parse the fully rendered HTML
tables = pd.read_html(driver.page_source)
driver.quit()

4. You have a critical syntax error in your code

Wait a second—your pandas line has a mistake: pd.read_html('url') uses the string literal 'url' instead of your url variable. That means pandas is trying to scrape a page at the invalid address http://url/, which obviously doesn’t exist. No wonder it’s throwing an error!

Fix: Remove the quotes around the variable:

tables = pd.read_html(url)  # Or better, use the response text from your logged-in session

5. Encoding issues are mangling the response

Sometimes servers return content with an unexpected encoding, causing requests.text to come out empty or garbled.

Fix: Force requests to auto-detect the correct encoding:

response = session.get('http://website/table')
# Let requests guess the right encoding
response.encoding = response.apparent_encoding
# Now check if response.text has content
print(response.text)

6. Your account lacks permission to access the table

Even after logging in, your account might not have the right permissions to view the table. Try manually logging into the site with your browser and navigating to the table URL—if you get a 403 error or an empty page there, that’s the issue.

Fix: Verify your account’s permissions, or check if you need to navigate through additional pages (instead of accessing the table directly via URL) to trigger the content load.


内容的提问来源于stack exchange,提问作者trob

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 07:29:08