Python新手爬取Wikifolio网站portfolio表格失败求助
Hey Julian, I totally get the frustration—spending hours testing 15+ methods and still hitting a wall with that portfolio data must be so annoying. Let’s break down why your current code isn’t working and fix it step by step.
Why Your Current Methods Fail
The core issue here is dynamic content loading. When you use requests.get() or urlopen() to fetch the page, you’re only getting the initial HTML sent by the server. The "Portfolio" table you want doesn’t exist in that initial HTML—it’s loaded after the page loads, via JavaScript pulling data from an internal API and rendering it on the page. That’s why your code keeps grabbing the static right-side table instead: it’s present from the start, while the portfolio rows aren’t.
Fix 1: Scrape the Direct API (Most Efficient)
Instead of wrestling with HTML parsing, you can pull the portfolio data directly from the API the website uses. This is faster and gives you clean, structured data.
Step 1: Find the API Endpoint
- Open the Wikifolio page in Chrome/Firefox, hit F12 to open DevTools.
- Go to the Network tab, refresh the page.
- Filter by "XHR" to narrow down requests—you’ll see a call that returns the portfolio data. For your page, it’s likely something like
https://www.wikifolio.com/api/de/de/instrument/wffalkinve/portfolio.
Step 2: Code to Fetch from the API
import requests import pandas as pd # Replace with the actual API endpoint you found in DevTools API_URL = "https://www.wikifolio.com/api/de/de/instrument/wffalkinve/portfolio" HEADERS = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36" } # Fetch the data response = requests.get(API_URL, headers=HEADERS) response.raise_for_status() # Raise error if request fails data = response.json() # Extract portfolio items and convert to DataFrame portfolio_items = data["data"]["items"] df = pd.DataFrame(portfolio_items) # Check key columns (adjust based on your needs) print(df[["name", "symbol", "share", "marketValue"]].head())
Fix 2: Use a Browser Renderer (If You Prefer HTML Parsing)
If you need to parse the fully rendered page (for example, if you want other elements too), use Selenium to simulate a browser that executes JavaScript.
Code with Selenium
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from bs4 import BeautifulSoup import pandas as pd URL = "https://www.wikifolio.com/de/de/w/wffalkinve" # Initialize Chrome driver (make sure ChromeDriver is installed and in your PATH) driver = webdriver.Chrome() driver.get(URL) try: # Wait up to 10 seconds for the portfolio rows to load WebDriverWait(driver, 10).until( EC.presence_of_element_located((By.CLASS_NAME, "c-portfolio")) ) # Get the fully rendered page source soup = BeautifulSoup(driver.page_source, "lxml") # Now find all the portfolio rows portfolio_rows = soup.find_all("tr", class_="c-portfolio") # Extract data from each row portfolio_data = [] for row in portfolio_rows: cells = row.find_all("td") row_content = [cell.get_text(strip=True) for cell in cells] portfolio_data.append(row_content) # Convert to DataFrame df = pd.DataFrame(portfolio_data) print(df.head()) finally: # Always close the browser driver.quit()
Notes for Selenium:
- You need to install the correct driver for your browser (e.g., ChromeDriver for Chrome, GeckoDriver for Firefox) and ensure it matches your browser version.
Quick Tips to Avoid Future Headaches
- Check for dynamic content first: If elements don’t show up in the initial HTML, they’re probably loaded via JS—use DevTools to find the underlying API.
- Respect
robots.txt: Always checkhttps://www.wikifolio.com/robots.txtto confirm scraping the portfolio data is allowed. - Use proper headers: Adding a User-Agent header helps avoid being blocked by the server.
内容的提问来源于stack exchange,提问作者Julian

