如何使用BeautifulSoup导入以JSON对象形式存在的CSS表格?
Hey there! Let's tackle your question step by step, plus fix a tiny issue in your current code first.
Quick Fix for Your Existing Code
I noticed a small syntax error in this line:
table_tag = tree.select(playersData)[0]
playersData should be a string selector (like a CSS ID or class). For example, if your table has an ID of playersData, it should be:
table_tag = tree.select('#playersData')[0] # Added quotes and # for ID selector
That should prevent a NameError from popping up.
How to Handle Tables Rendered from JSON Data
BeautifulSoup doesn't execute JavaScript, so if the table is built by JS using JSON data after the page loads, we need to adjust our approach based on where the JSON lives:
1. JSON is Embedded in a <script> Tag in the HTML
Many sites embed JSON directly in the page's <script> tags (often with type="application/json"). We can extract this script content, parse it as JSON, then convert it to a CSV table.
Here's a working example:
import csv import json from bs4 import BeautifulSoup import urllib.request as ur # Set up CSV writer with UTF-8 encoding (prevents weird character issues) outfile = open(r"table_data.csv", "w+", newline='', encoding='utf-8') writer = csv.writer(outfile) html = ur.urlopen('your_target_url_here') tree = BeautifulSoup(html, "lxml") # Find the script tag containing your JSON data (adjust the selector to match your page) script_tag = tree.find('script', id='tableData') # Example: script with ID "tableData" # Parse the JSON content from the script tag json_data = json.loads(script_tag.string.strip()) # Extract the table data (adjust the key to match your JSON structure) table_rows = json_data.get('players', []) if table_rows: # Write the header row using the keys from the first JSON object writer.writerow(table_rows[0].keys()) # Write each data row for row in table_rows: writer.writerow(row.values()) print(' '.join(map(str, row.values()))) outfile.close()
2. JSON is Loaded via an AJAX Request
If the JSON comes from a separate API call (you'll see this in your browser's Network tab under XHR/Fetch requests), you don't even need BeautifulSoup—just request the API endpoint directly:
import csv import json import urllib.request as ur outfile = open(r"table_data.csv", "w+", newline='', encoding='utf-8') writer = csv.writer(outfile) # Replace this with the actual API URL you found in DevTools api_url = "https://example.com/api/your-table-data" response = ur.urlopen(api_url) # Decode the response and parse as JSON json_data = json.loads(response.read().decode('utf-8')) table_rows = json_data.get('players', []) if table_rows: writer.writerow(table_rows[0].keys()) for row in table_rows: writer.writerow(row.values()) print(' '.join(map(str, row.values()))) outfile.close()
To find the API URL:
- Open your browser's DevTools (F12)
- Go to the Network tab
- Refresh the page
- Look for requests labeled XHR/Fetch that return JSON data matching your table content
内容的提问来源于stack exchange,提问作者Pranav Pandit

