关于使用Supervised Learning Linear Regression构建桌游评估函数的疑问
Hey Cong, this is such a relatable confusion when trying to bridge traditional game AI with machine learning—let’s break this down clearly so you can move past the contradiction you’re feeling.
Your core sticking point is: "I need labeled output data (state evaluation values) to train linear regression, but writing a manual evaluation function feels like defeating the purpose of using linear regression to generate one." The good news is you don’t have to write a full, polished manual evaluation function to get usable labels. Here are three practical, non-contradictory approaches:
1. Use real game outcome data as labels (most straightforward)
Instead of manually assigning scores to states, leverage actual gameplay results from skilled players. For each board state in your collected game logs:
- Label it with 1 if the player whose turn it was eventually won the game
- Label it with 0 if they lost
- For more granularity, you can use the probability of winning from that state (e.g., if a state appears in 100 games and the current player wins 65 of them, label it 0.65)
This approach requires no manual evaluation logic—you’re letting the actual results of human play define what a "good" or "bad" state is. Linear regression will then learn to map board state features (like piece counts, position metrics, control of key areas) to these outcome-based labels.
2. Generate labels via Monte Carlo Tree Search (MCTS) simulations
If you don’t have access to enough human game data, use MCTS to simulate thousands of games from each board state. The label for a state becomes the win rate of the current player across all those simulations.
MCTS doesn’t rely on a pre-written evaluation function—it uses random playouts to estimate state value. This means you’re letting the game’s rules themselves generate the labels, and linear regression will learn to generalize those simulation results into a fast, lightweight evaluation function.
3. Weak labels + iterative refinement (middle ground)
If you do have some basic, simple heuristics (e.g., "having more high-value pieces = better state"), you can use these to generate initial "weak labels" for a small set of states. Then:
- Train your linear regression model on these weak labels
- Use the model to predict labels for a larger set of unlabeled states
- Validate these predictions against MCTS simulations or human game outcomes, and refine the training data
This doesn’t defeat your goal—your manual heuristics are just a starting point, not the final evaluation function. The linear regression model will learn to improve on those heuristics over time, incorporating patterns you might not have noticed manually.
Key Takeaway
You never need to write a complete, final evaluation function upfront. The labels for your linear regression model can come directly from game outcomes or automated simulations, which aligns perfectly with your goal of using ML to generate the evaluation function itself.
内容的提问来源于stack exchange,提问作者Cong Yang

