使用MechanicalSoup访问隐藏表单时触发“无闭合引号”值错误
Hey there! I totally get how frustrating it is when basic login works but more complex form operations throw weird parsing errors—let's break down what's going on and how to fix it.
Why This Happens
That error almost always boils down to malformed HTML in the hidden form fields. Websites sometimes generate messy HTML for hidden inputs (like missing closing quotes around attribute values, or unescaped special characters in values), and MechanicalSoup's underlying parser (usually lxml) is strict about following proper HTML standards. Your Chrome dev tools are way more forgiving with wonky HTML, which is why it worked there but not in MechanicalSoup.
Fixes to Try
1. Switch to a More Forgiving HTML Parser
MechanicalSoup uses BeautifulSoup under the hood, and you can tell it to use the built-in html.parser instead of lxml—this parser is more lenient with malformed HTML. Update your browser initialization code like this:
import mechanicalsoup # Use html.parser instead of the default lxml browser = mechanicalsoup.StatefulBrowser(soup_config={'features': 'html.parser'})
This often fixes quotation-related parsing errors right away.
2. Manually Correct or Set Hidden Field Values
If switching parsers doesn't work, you can bypass automatic form parsing entirely by manually defining the hidden fields you need. First, use your browser's dev tools to grab the exact names and values of all hidden fields in the form, then construct the POST data yourself:
# Skip selecting the form directly—build the data manually form_data = { # Your login credentials (if needed) "login[username]": "your_email@example.com", "login[password]": "your_password", # Add all hidden fields you found in dev tools "form_key": "abc123xyz", "return_url": "https://www.thegoodwillout.de/", # Any other hidden inputs from the form } # Send the POST request directly response = browser.post( "https://www.thegoodwillout.de/customer/account/loginPost/", data=form_data )
This avoids parsing the problematic HTML altogether by using the exact data the browser would send.
3. Manually Clean Up the HTML Before Parsing
If you need to use the form object for other operations, you can fetch the raw HTML, fix the quotation issues with a regex, then feed it to BeautifulSoup manually:
# Fetch the raw HTML page_html = browser.get("your_form_page_url").text # Fix unquoted attribute values (example regex—adjust as needed) import re cleaned_html = re.sub(r'(value=)([^\s">]+)', r'\1"\2"', page_html) # Parse the cleaned HTML with BeautifulSoup from bs4 import BeautifulSoup soup = BeautifulSoup(cleaned_html, "html.parser") # Now select the form from the cleaned soup form = mechanicalsoup.Form(soup.select_one("form[action*='loginPost']")) browser.form = form
This regex finds instances like value=somevalue and wraps the value in quotes to make it valid HTML.
Final Notes
If none of these work, double-check the hidden fields in your browser's Network tab—sometimes websites dynamically inject hidden fields via JavaScript, which MechanicalSoup (since it doesn't run JS) won't see. In that case, you might need to use a tool that can execute JavaScript, but that's a bigger jump.
Hope one of these fixes gets you past that error!
内容的提问来源于stack exchange,提问作者daehruoydeef

