从HTML提取的JSON字符串验证异常:JSONLint有效代码验证无效
It’s super frustrating when JSONLint says your string is valid, but your code throws a parsing error—here are the most common culprits and fixes:
Hidden Unicode Gremlins: HTML often sneaks in invisible characters like non-breaking spaces (
\u00A0), zero-width spaces, or curly quotes that JSONLint automatically cleans up, but your parser doesn’t. For example, a non-breaking space in a string value will break most JSON parsers even though JSONLint ignores it.Truncated JSON: Your sample shows the JSON is cut off (
thumb":"https://www...). If you’re parsing this incomplete snippet, even a partial valid chunk will fail in code—make sure you’re extracting the full JSON array/object from the HTML.Unencoded HTML Entities: Sometimes JSON in HTML is stored with escaped entities like
"instead of", or<for angle brackets. JSONLint decodes these on the fly, but your code might not be handling this step. For example, if the HTML has{"shortid":"456673",...}, you need to decode those entities before parsing.Extra Junk Around the JSON: HTML might have leading/trailing whitespace, newlines, or even HTML comments wrapped around the JSON. JSONLint ignores this extra stuff, but strict parsers will choke if there’s anything outside the root
{}or[]. Also, JSON doesn’t allow comments—if the HTML has//or/* */inside the JSON, JSONLint strips them, but your code will throw an error.
Quick Fixes to Try
Clean the String First: Strip hidden characters and whitespace before parsing.
In JavaScript:const cleanedJson = extractedString.replace(/[\u00A0\u200B]/g, '').trim();In Python:
import re cleaned_json = re.sub(r'[\u00A0\u200B]', '', extracted_string).strip()Decode HTML Entities: Use a library to handle entity decoding. For JavaScript, use the
hepackage’she.decode()method, orDOMParserin the browser. In Python, the standard library’shtml.unescape()works great.Double-Check Extraction: Make sure you’re capturing the entire JSON block. If it’s inside a
<script>tag, use a regex that grabs everything from the opening[/{to the closing]/}—avoid truncating it mid-value.Debug the Raw String: Log the exact string your code is trying to parse. Use string escaping to see hidden Unicode characters (like
console.log(JSON.stringify(extractedString))in JS) to spot any weirdness you can’t see with the naked eye.
Pro Tip: If possible, skip extracting JSON from HTML entirely. Check if the page loads the data via an API endpoint—scraping that directly is way more reliable and avoids these parsing headaches.
内容的提问来源于stack exchange,提问作者Im Batman

