JustDial数据爬取Python代码报错json.decoder.JSONDecodeError求助
Hey there, let's sort out that JSONDecodeError you're running into with your JustDial scraping code. I've looked over your script and found a couple of key issues that are causing this problem—here's a step-by-step guide to get it working:
Step-by-Step Modifications
1. Fix the Broken URL Concatenation
Your URL was split across two lines without proper concatenation, which meant the final request was using an incomplete address. Additionally, the & entities are HTML-specific; we need to replace them with actual & characters for the URL to work correctly.
Replace your original URL line with this:
url = 'http://www.justdial.com/functions/ajxsearch.php?national_search=0&act=pagination&city={0}&search={1}&page={2}'
2. Add Browser-like Request Headers
JustDial actively blocks non-browser requests, which is probably why you're getting non-JSON responses (like a captcha page) that crash the JSON parser. We'll add headers to mimic a real Chrome browser:
Add this right after your imports:
# Mimic a browser request to avoid being blocked headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36', 'Referer': 'http://www.justdial.com/' }
Then update your request line to include these headers:
resp = requests.get(url.format(where, what, page), headers=headers)
3. Add Error Handling for JSON Parsing
Even with headers, you might occasionally get non-JSON responses. We'll wrap the JSON parsing in a try-except block to handle this gracefully instead of crashing:
Replace the lines starting with if not resp.json()['paidDocIds']: with this:
# Try to parse the JSON response, handle failures try: resp_json = resp.json() except json.JSONDecodeError: print(f"⚠️ Failed to parse JSON on page {page} — this might be a block or broken response.") page += 1 continue # Check if we've reached the end of results if 'paidDocIds' not in resp_json or not resp_json['paidDocIds']: print(25*"-") print('Scraping Finished') print(25*"-") break
4. Simplify and Stabilize Data Extraction
Your original data extraction logic uses multiple loops, which can be inefficient and prone to missing data. Let's simplify it to directly target the elements we need, and add fallbacks for missing data to keep your CSV columns aligned:
Replace the for thing in soup.find_all('section'): block with this:
# Directly target sections with the 'jcar' class for thing in soup.find_all('section', class_='jcar'): csv_list = [] # Extract company name company_tag = thing.find('a', onclick="_ct('clntnm', 'lspg');") csv_list.append(company_tag.get('title') if company_tag else "") # Extract address address_tag = thing.find('span', class_=['mrehover', 'dn']) csv_list.append(address_tag.get_text().strip() if address_tag else "") # Extract phone number phone_tag = thing.find('a', href=lambda x: x and x.startswith('tel:')) csv_list.append(phone_tag.get('href').split(':')[-1] if phone_tag else "") # Escape quotes for valid CSV formatting csv_list = ['"{}"'.format(item.replace('"', '""')) for item in csv_list] f.write(','.join(csv_list) + '\n')
Full Modified Code
Here's the complete working script with all the above changes:
import csv import json import requests from bs4 import BeautifulSoup # Mimic a browser request to avoid being blocked headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36', 'Referer': 'http://www.justdial.com/' } print(25*"=") print("Just Dial Scraper") print(25*"=") url = 'http://www.justdial.com/functions/ajxsearch.php?national_search=0&act=pagination&city={0}&search={1}&page={2}' what = input("Enter your Query: ") what = what.replace(' ', '+') where = input("Enter the Location: ") with open(f"{what}_{where}.csv", 'w', newline='', encoding='utf-8') as f: writer = csv.writer(f) writer.writerow(['company', 'address', 'phone']) page = 1 while True: print('Scraping Page', page) resp = requests.get(url.format(where, what, page), headers=headers) # Try to parse the JSON response, handle failures try: resp_json = resp.json() except json.JSONDecodeError: print(f"⚠️ Failed to parse JSON on page {page} — this might be a block or broken response.") page += 1 continue # Check if we've reached the end of results if 'paidDocIds' not in resp_json or not resp_json['paidDocIds']: print(25*"-") print('Scraping Finished') print(25*"-") break markup = resp_json['markup'].replace('\/', '/') soup = BeautifulSoup(markup, 'html.parser') # Directly target sections with the 'jcar' class for thing in soup.find_all('section', class_='jcar'): csv_list = [] # Extract company name company_tag = thing.find('a', onclick="_ct('clntnm', 'lspg');") csv_list.append(company_tag.get('title') if company_tag else "") # Extract address address_tag = thing.find('span', class_=['mrehover', 'dn']) csv_list.append(address_tag.get_text().strip() if address_tag else "") # Extract phone number phone_tag = thing.find('a', href=lambda x: x and x.startswith('tel:')) csv_list.append(phone_tag.get('href').split(':')[-1] if phone_tag else "") writer.writerow(csv_list) page += 1
Note: I also switched to using Python's built-in csv.writer instead of manual string concatenation—it's more reliable for handling edge cases like commas or quotes in the data.
内容的提问来源于stack exchange,提问作者M. Patel

