You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

JustDial数据爬取Python代码报错json.decoder.JSONDecodeError求助

Fixing the JSONDecodeError in Your JustDial Scraper

Hey there, let's sort out that JSONDecodeError you're running into with your JustDial scraping code. I've looked over your script and found a couple of key issues that are causing this problem—here's a step-by-step guide to get it working:

Step-by-Step Modifications

1. Fix the Broken URL Concatenation

Your URL was split across two lines without proper concatenation, which meant the final request was using an incomplete address. Additionally, the & entities are HTML-specific; we need to replace them with actual & characters for the URL to work correctly.

Replace your original URL line with this:

url = 'http://www.justdial.com/functions/ajxsearch.php?national_search=0&act=pagination&city={0}&search={1}&page={2}'

2. Add Browser-like Request Headers

JustDial actively blocks non-browser requests, which is probably why you're getting non-JSON responses (like a captcha page) that crash the JSON parser. We'll add headers to mimic a real Chrome browser:

Add this right after your imports:

# Mimic a browser request to avoid being blocked
headers = {
    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36',
    'Referer': 'http://www.justdial.com/'
}

Then update your request line to include these headers:

resp = requests.get(url.format(where, what, page), headers=headers)

3. Add Error Handling for JSON Parsing

Even with headers, you might occasionally get non-JSON responses. We'll wrap the JSON parsing in a try-except block to handle this gracefully instead of crashing:

Replace the lines starting with if not resp.json()['paidDocIds']: with this:

# Try to parse the JSON response, handle failures
try:
    resp_json = resp.json()
except json.JSONDecodeError:
    print(f"⚠️ Failed to parse JSON on page {page} — this might be a block or broken response.")
    page += 1
    continue

# Check if we've reached the end of results
if 'paidDocIds' not in resp_json or not resp_json['paidDocIds']:
    print(25*"-")
    print('Scraping Finished')
    print(25*"-")
    break

4. Simplify and Stabilize Data Extraction

Your original data extraction logic uses multiple loops, which can be inefficient and prone to missing data. Let's simplify it to directly target the elements we need, and add fallbacks for missing data to keep your CSV columns aligned:

Replace the for thing in soup.find_all('section'): block with this:

# Directly target sections with the 'jcar' class
for thing in soup.find_all('section', class_='jcar'):
    csv_list = []
    
    # Extract company name
    company_tag = thing.find('a', onclick="_ct('clntnm', 'lspg');")
    csv_list.append(company_tag.get('title') if company_tag else "")
    
    # Extract address
    address_tag = thing.find('span', class_=['mrehover', 'dn'])
    csv_list.append(address_tag.get_text().strip() if address_tag else "")
    
    # Extract phone number
    phone_tag = thing.find('a', href=lambda x: x and x.startswith('tel:'))
    csv_list.append(phone_tag.get('href').split(':')[-1] if phone_tag else "")
    
    # Escape quotes for valid CSV formatting
    csv_list = ['"{}"'.format(item.replace('"', '""')) for item in csv_list]
    f.write(','.join(csv_list) + '\n')

Full Modified Code

Here's the complete working script with all the above changes:

import csv
import json
import requests
from bs4 import BeautifulSoup

# Mimic a browser request to avoid being blocked
headers = {
    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36',
    'Referer': 'http://www.justdial.com/'
}

print(25*"=")
print("Just Dial Scraper")
print(25*"=")
url = 'http://www.justdial.com/functions/ajxsearch.php?national_search=0&act=pagination&city={0}&search={1}&page={2}'
what = input("Enter your Query: ")
what = what.replace(' ', '+')
where = input("Enter the Location: ")

with open(f"{what}_{where}.csv", 'w', newline='', encoding='utf-8') as f:
    writer = csv.writer(f)
    writer.writerow(['company', 'address', 'phone'])
    
    page = 1
    while True:
        print('Scraping Page', page)
        resp = requests.get(url.format(where, what, page), headers=headers)
        
        # Try to parse the JSON response, handle failures
        try:
            resp_json = resp.json()
        except json.JSONDecodeError:
            print(f"⚠️ Failed to parse JSON on page {page} — this might be a block or broken response.")
            page += 1
            continue
        
        # Check if we've reached the end of results
        if 'paidDocIds' not in resp_json or not resp_json['paidDocIds']:
            print(25*"-")
            print('Scraping Finished')
            print(25*"-")
            break
        
        markup = resp_json['markup'].replace('\/', '/')
        soup = BeautifulSoup(markup, 'html.parser')
        
        # Directly target sections with the 'jcar' class
        for thing in soup.find_all('section', class_='jcar'):
            csv_list = []
            
            # Extract company name
            company_tag = thing.find('a', onclick="_ct('clntnm', 'lspg');")
            csv_list.append(company_tag.get('title') if company_tag else "")
            
            # Extract address
            address_tag = thing.find('span', class_=['mrehover', 'dn'])
            csv_list.append(address_tag.get_text().strip() if address_tag else "")
            
            # Extract phone number
            phone_tag = thing.find('a', href=lambda x: x and x.startswith('tel:'))
            csv_list.append(phone_tag.get('href').split(':')[-1] if phone_tag else "")
            
            writer.writerow(csv_list)
        
        page += 1

Note: I also switched to using Python's built-in csv.writer instead of manual string concatenation—it's more reliable for handling edge cases like commas or quotes in the data.

内容的提问来源于stack exchange,提问作者M. Patel

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.28 10:16:12