如何为Boost.Spirit.Qi列表运算符(%)优化内存分配解析超大列表?
Absolutely! You can avoid those expensive log₂(N) memory reallocations by pre-allocating your vector using the initial size value from the input. Boost.Spirit.Qi offers a couple of straightforward ways to do this, depending on whether you just want to optimize memory usage or also enforce strict validation of the element count.
Approach 1: Simple Capacity Reservation (For Trusted Input)
If you’re confident the input will always have exactly N elements matching the initial size, you can reserve the vector’s capacity upfront before parsing the list. This ensures the vector has enough space to hold all elements without needing to reallocate during parsing.
Here’s how to implement it using Phoenix semantic actions to call std::vector::reserve:
#include <boost/spirit/include/qi.hpp> #include <boost/spirit/include/phoenix_core.hpp> #include <boost/spirit/include/phoenix_stl.hpp> #include <vector> #include <string> namespace qi = boost::spirit::qi; namespace phx = boost::phoenix; int main() { std::string input = "100000; 1, 2, 3, ..., 100000;"; // Your large input std::vector<int> data; bool success = qi::parse( input.begin(), input.end(), // Parse initial size, reserve capacity, then parse the list qi::int_[phx::reserve(qi::_val, qi::_1)] >> qi::lit(';') >> qi::int_ % qi::lit(',') >> qi::lit(';'), data ); if (success) { // Use your pre-allocated vector without reallocations } return 0; }
How It Works:
- The first
qi::int_parses the initial sizeN. - The semantic action
phx::reserve(qi::_val, qi::_1)callsdata.reserve(N), setting the vector’s capacity to at leastN. - The
qi::int_ % ','parses the list of elements, appending them to the vector. Since capacity is already sufficient, no reallocations occur.
Approach 2: Strict Element Count Validation (For Un-Trusted Input)
If you need to ensure the input exactly matches the declared size N (failing if there are too few or too many elements), you can resize the vector upfront and parse each element directly into its pre-allocated position using the repeat iterator placeholder qi::_i.
#include <boost/spirit/include/qi.hpp> #include <boost/spirit/include/phoenix_core.hpp> #include <boost/spirit/include/phoenix_operator.hpp> #include <vector> #include <string> namespace qi = boost::spirit::qi; namespace phx = boost::phoenix; template <typename Iterator> struct StrictListParser : qi::grammar<Iterator, std::vector<int>()> { StrictListParser() : StrictListParser::base_type(start) { start = // Parse size and resize vector to hold exactly N elements qi::int_[phx::resize(qi::_val, qi::_1)] >> qi::lit(';') // Parse exactly N elements, assigning each to its pre-allocated index >> qi::repeat(qi::_1)[ qi::int_[qi::_val[qi::_i] = qi::_1] >> -(qi::lit(',')) // Allow optional trailing comma (adjust if needed) ] >> qi::lit(';'); } qi::rule<Iterator, std::vector<int>()> start; }; int main() { std::string input = "5; 1,2,3,4,5;"; std::vector<int> data; StrictListParser<std::string::iterator> parser; bool success = qi::parse(input.begin(), input.end(), parser, data); if (success) { // data has exactly 5 elements, no reallocations occurred } return 0; }
How It Works:
phx::resize(qi::_val, qi::_1)sets the vector’s size toN, creating default-initialized elements.qi::repeat(qi::_1)ensures exactlyNelements are parsed (the parser will fail if the count doesn’t match).qi::_iis a placeholder for the current repetition index (0-based), so each parsed int is assigned directly to the corresponding position in the pre-resized vector.- The
-(qi::lit(','))handles optional commas between elements (remove this part if your input never has trailing commas).
Key Takeaways:
- Both approaches eliminate the
log₂(N)reallocations by pre-allocating memory upfront. - The first approach is simpler for trusted inputs, while the second adds validation to catch mismatches between the declared size and actual element count.
Content of the question originates from Stack Exchange, question author: Riyaaaaa

