使用browserless.js获取的URL中Unicode转义无法解码的技术求助
u003d) in URLs Scraped with Browserless.js I’ve dealt with similar weird encoding glitches when scraping with headless browsers before, so let’s unpack what’s going on here and get those images loading properly.
The Root Cause
Your hint about matching u003d (not \u003d) working tells me everything: the URLs you’re getting don’t contain actual Unicode-escaped = characters (where \u003d resolves to =), but instead literal text sequences of u003d. Browserless must be transforming original = characters into this literal string format during DOM serialization—something that doesn’t happen when you copy the URL directly from the console (your browser probably auto-converts the literal u003d back to = when you paste it elsewhere).
Quick Fixes
1. Target the Literal u003d Directly
Since you confirmed matching u003d works, just replace those sequences with actual = signs:
// Replace all literal u003d instances with = const fixedImageUrl = scrapedUrl.replace(/u003d/g, '=');
This should immediately fix the broken image issue for your example URL.
2. Handle All Literal u00XX Escapes (For Future-Proofing)
If you run into other similar escapes (like u0026 for &), use a more flexible regex to convert any literal u00XX sequence to its corresponding ASCII character:
const fixedImageUrl = scrapedUrl.replace(/u00([0-9A-Fa-f]{2})/g, (match, hexCode) => { // Convert the 2-digit hex code to its ASCII character return String.fromCharCode(parseInt(hexCode, 16)); });
This will handle any standard ASCII character that’s been converted to a literal u00XX string.
Why Your Previous Methods Didn’t Work
Let’s clear up why the approaches you tried failed:
JSON.parse(JSON.stringify(...)): This only resolves actual JSON escape sequences (like\u003dinside a JSON string). Since your string has literalu003d(not escaped), JSON doesn’t recognize it as a special sequence.String.prototype.normalize(): This is for Unicode normalization (e.g., combining accent characters), which has nothing to do with ASCII escape strings.decodeURIComponent(): This decodes URI-encoded characters like%3D, not literalu003dtext.
Verify the Issue (Optional)
To confirm this is indeed the problem, run a quick check on your scraped URL string:
console.log(scrapedUrl.includes('u003d')); // Should return true console.log(scrapedUrl.includes('\u003d')); // Should return false (since \u003d is just =)
If those results match, you know you’re dealing with literal text sequences rather than actual Unicode escapes.
内容的提问来源于stack exchange,提问作者user11909477

