解决Python Web代理的DNSLookupFailedError问题
Hey there! Let's work through this proxy resource matching issue you're facing. It sounds like your transform_content.py is missing patterns for certain resource URLs, causing errors when accessing specific sites. Here's a breakdown of steps to fix this, plus alternative library options if you want to simplify things:
Before tweaking code, figure out exactly which URLs aren't being caught:
- Open your browser's DevTools > Network tab when accessing the problematic site.
- Look for requests that return 404 errors, or that load directly from the original domain instead of your proxy.
- Note the URL formats: are they protocol-relative (like
//example.com/script.js), relative paths with../, special resource types (.woff2fonts,.webpimages), or hidden in attributes likedata-src/srcset?
transform_content.py Chances are your existing regex or string matching isn't covering edge cases. Common gaps include:
- Missing support for single-quoted attributes (e.g.,
src='image.png'instead ofsrc="image.png"). - Ignoring CSS
background: url('/bg.jpg')style declarations. - Not handling
srcsetattributes for responsive images. - Skipping protocol-relative URLs or root-relative paths.
For example, if your current code uses a regex like r'src="([^"]+)"', it won't catch single-quoted values or unquoted attributes. You can expand it to cover more cases:
# Match src/href with single quotes, double quotes, or no quotes content = re.sub(r'(src|href)=["\']?([^"\'>\s]+)["\']?', lambda m: f'{m.group(1)}="{proxy_url(m.group(2))}"', content)
Here are specific patterns to add to your transformation logic:
- CSS background URLs:
# Replace url() values in inline styles and style tags content = re.sub(r'url\((["\']?)([^"\')]+)\1\)', lambda m: f'url("{proxy_url(m.group(2))}")', content) - Responsive image srcset:
# Split srcset values and proxy each URL def proxy_srcset(srcset): return ', '.join([f'{proxy_url(url)} {size}' if size else proxy_url(url) for url, size in (s.strip().split(' ', 1) for s in srcset.split(','))]) content = re.sub(r'srcset=["\']?([^"\'>]+)["\']?', lambda m: f'srcset="{proxy_srcset(m.group(1))}"', content) - Data attributes like data-src:
content = re.sub(r'data-src=["\']?([^"\'>\s]+)["\']?', lambda m: f'data-src="{proxy_url(m.group(2))}"', content)
If maintaining regex patterns feels brittle, use dedicated HTML parsing libraries to reliably traverse and modify content:
Option 1: BeautifulSoup
This library makes it easy to iterate over all HTML tags and attributes without regex headaches:
from bs4 import BeautifulSoup def transform_content(content, base_url, proxy_prefix): soup = BeautifulSoup(content, 'html.parser') # Process standard resource tags for tag in soup.find_all(['img', 'script', 'link', 'a', 'source']): for attr in ['src', 'href', 'srcset']: if tag.get(attr): tag[attr] = proxy_url(tag[attr], base_url, proxy_prefix) # Process inline styles and style tags for style_tag in soup.find_all('style'): if style_tag.string: style_tag.string = re.sub(r'url\((["\']?)([^"\')]+)\1\)', lambda m: f'url("{proxy_url(m.group(2), base_url, proxy_prefix)}")', style_tag.string) for tag in soup.find_all(style=True): tag['style'] = re.sub(r'url\((["\']?)([^"\')]+)\1\)', lambda m: f'url("{proxy_url(m.group(2), base_url, proxy_prefix)}")', tag['style']) return str(soup)
Option 2: lxml
For faster parsing of large content, lxml is a great alternative (works similarly to BeautifulSoup but with better performance):
from lxml import html def transform_content(content, base_url, proxy_prefix): tree = html.fromstring(content) # Process attributes for attr in ['src', 'href', 'srcset']: for element in tree.xpath(f'//*[@{attr}]'): original_url = element.get(attr) element.set(attr, proxy_url(original_url, base_url, proxy_prefix)) # Process CSS styles for style_element in tree.xpath('//style'): if style_element.text: style_element.text = re.sub(r'url\((["\']?)([^"\')]+)\1\)', lambda m: f'url("{proxy_url(m.group(2), base_url, proxy_prefix)}")', style_element.text) return html.tostring(tree, encoding='unicode')
Since you've added a test suite, create test cases that cover the problematic URL patterns you identified. For example:
- Test protocol-relative URLs (
//example.com/style.css) - Test relative paths with
../(../assets/logo.png) - Test CSS background URLs and
srcsetattributes
Run the tests after each change to ensure your fixes don't break existing functionality.
If you can share specific error logs or examples of URLs that aren't being proxied, we can narrow down the exact pattern to add!
内容的提问来源于stack exchange,提问作者T145

