My NHL Lineup Optimizer Now Skips the Solver (Mostly)

My NHL lineup optimizer used to build every lineup the same way: hand the whole problem to a solver and wait. It now skips the solver for most of the work. It lists the stacks that are allowed, scores them, and picks the best one, and it does that for a whole batch on a DraftKings or FanDuel slate in about a tenth of a second. The solver I have always reached for is still in the tool, but as the referee rather than the player.

Horizontal bar chart on a log scale. Time to build a batch of lineups with the solver versus the enumeration fast path: DraftKings 3-3 stacks 2.3 versus 0.1 seconds, FanDuel 3-2 stacks with cooldown off 44.9 versus 0.15 seconds.
Time to build the whole batch. The 3-3 rows were run through the app code on tonight’s DraftKings and FanDuel slates; the 3-2 rows were run against the engine directly on a FanDuel slate. FanDuel leaves its own projections blank this early in the season, so I used stand-ins, which do not change the timing. In every row the two methods produced the same lineup totals.

Why the solver was slow

A lineup optimizer does not solve one problem, it solves a chain of them. Lineup 1 is the best lineup. Lineup 2 is the best lineup that is different enough from lineup 1. Lineup 60 has to be different enough from all 59 before it, and every one of those differences is another constraint. On the NHL model that came to about two seconds per lineup with CBC, the free solver I use, and the cost was never in the salary cap or the positions. It was in choosing the stacks.

I spent a fair amount of time trying to speed that up and most of it failed. Shrinking the player pool made it slower. A looser solver gap made it slower or no better. HiGHS, which is the right answer for my staff scheduler, was twice as slow as CBC here. A textbook tightening of the model made it twice as slow. The lesson I kept relearning is that the shape of the model decides, not the solver’s reputation.

The trick: list the stacks instead of solving for them

Here is what I noticed. The hard part of the problem is a choice among a small number of things. A hockey team has four forward lines of three players, and the stacking rules I use say things like “one full line, plus a second line from another team”. A six-team slate has 24 forward lines, and with the default setting of using only lines 1 and 2, it has 12. Choosing two of them from different teams gives 60 combinations. That is not a problem for a solver. That is a list.

So the fast path does this, in order:

  • List every allowed stack. For the 3-3 setup that is a pair of lines on different teams; for the 3-2 setup it is one full line plus two or three players from a line on another team.
  • Give each one an optimistic score. That is the stack’s projected points plus the best the rest of the roster could possibly add: the best goalie, the best defence pair that fits the remaining salary, and the best remaining forwards. It is an upper bound, and it is never lower than the truth.
  • Take them best first. Work out the real best lineup for the most promising stack, including every rule about earlier lineups. Keep it, then move on. Stop as soon as the next stack’s optimistic score cannot beat what you already have.
  • Fill the roster. The skaters that are not part of the stack are chosen with array operations over defence pairs, rather than a solver.

Because the bound is honest, the first lineup that beats every remaining bound is the best one, not just a good one. This is the same idea as branch and bound inside a solver, with one difference: I am choosing what to branch on, and I am choosing something that hockey gives me for free.

Why hockey suits it

Enumeration only works when the thing you are choosing is small and discrete. Hockey has that. Lines are fixed trios that play together, pairings are fixed, and a slate has a couple of dozen of each. The scoring is tied to those units, since a line scores together. Compare that with a sport where a “stack” is any three of nine batters in a window of the batting order, or a quarterback with whichever receivers you like. Those have far more ways to build a stack, and a search like this would drown.

It is also why the rules I use are easy to check exactly. Every earlier lineup can be tested directly against a candidate: how many players they share, whether the top three or five by salary repeat, whether a line is coming back too soon. A solver has to write each of those as a constraint. The fast path just counts.

The solver is still the referee

I am an operations research person, and I will admit it was a little disappointing. After all the modelling, the fastest thing turned out not to be math programming at all. It is a sort with some careful bookkeeping. The consolation is that I still have a mathematical programming model of the same problem, and it can grade the fast path’s homework.

That is what I did. In testing, every time the fast path produced a lineup, I asked the solver for the best lineup given the same earlier lineups and compared the totals. I ran this across both sites, both stack setups, and a batch of randomised pools. If the two ever disagreed, the fast path was wrong and I would rather find out in a test than in a contest.

The bug the referee almost missed

The step-by-step check passed everywhere. Then I ran whole batches through the actual app on tonight’s FanDuel slate and compared the two methods end to end, and they disagreed: lineup 4 scored 99.4 with the fast path and 98.2 with the solver. Both were following the rules. They had built the first three lineups identically, yet chose different fourth lineups.

The cause was a tie. One of my diversity rules looks at the top three, five and seven players of each earlier lineup by salary, and caps how many of them can come back. FanDuel has lots of players at the same price, so “top three by salary” depends on the order the players are listed in, and the two methods list them differently. My step check could not see this, because it re-solved from the fast path’s own earlier lineups, so it was only checking the fast path against itself.

The fix is one line in each place: break salary ties by player ID. After that, every batch I compared matched. The moral is the usual one for testing: a check that shares your assumptions will pass your bugs. I found this one only because I also compared whole runs.

DraftKings 150, FanDuel 40, Yahoo 6

The number of lineups you can enter in one contest decides how much this matters. As I understand it, the maximum is 150 on DraftKings, 40 on FanDuel and 6 on Yahoo. At 150 lineups the solver’s two seconds each is about five minutes, and on a hosted server with less than one CPU it does not finish in a web request. At 40 it is a minute or two. At 6 it is six solves, which I do not think justifies a second engine, so I left Yahoo on the solver. That is my judgement, not a measurement I made on Yahoo.

What it does not cover

The fast path handles what I use most: 3-3 and 3-2 line stacks on DraftKings and FanDuel, with the usual diversity rules, and locked goalies. Everything else falls back to CBC without anyone noticing, except in the time it takes:

  • Locked skaters together with 3-2 stacks.
  • Power-play unit stacking.
  • The Cash preset, Vegas-based rules and strategy mixes.
  • The goalie overlap cap, the score slope, and any change to the limit on teams with defencemen.
  • Yahoo.

There is one more case. When the fast path runs out of lineups that meet the rules, it says so and the solver is asked to confirm. If the solver finds one the fast path missed, it is used; if it says nothing is left, the batch stops. So the solver does both the hard cases and the last check, and any error in the fast path also sends the work back to the solver. There is a switch on the server to turn the fast path off completely. I would rather it fall back a little too often than be wrong even once.

I am also fine with it only handling most of a batch. A tool that gets the common case in a tenth of a second and the rare case in the old time is a big improvement over one that does everything in the old time.

What the numbers do and do not say

  • Small slates are limited by the rules, not the engine. Tonight’s slate has three games. With the default unit cooldown, which keeps a line from being stacked again within nine lineups, both methods stop at 8 lineups on DraftKings. Turn the cooldown off and the fast path got to 38 before it ran out. Speed does not create lineups the rules do not allow.
  • The timings are from a test machine, not the live site. The hosted tool runs on a small server and is slower, as I wrote in the post about the MIT paper, so the ratio matters more than the seconds.
  • Same optimum does not mean winning lineups. The fast path reproduces the solver’s answer, nothing more. Whether those lineups win contests is a separate question that this post does not try to answer.

The 3-2 stacking setup comes from Picking Winners in Daily Fantasy Sports Using Integer Programming, the MIT paper behind my goalie build. They solved their version with integer programming, and so did I. For my version of this problem, with lines and pairings as the units, a search was the better tool, with the integer program as the referee. If you want the solver behind it, it is CBC from COIN-OR.

You can try the result in the free NHL lineup optimizer. Nothing about how you use it changes; it is just quicker on DraftKings and FanDuel.

Tradeline Supply
Things that I use, like, and am affiliated with:
Mint Mobile offers great cell phone service for $15 flat, get $15 off using the link. Get discounted phones with service activation and no contract.
I never spend money before I check Mr Rebates or Rakuten to get cashbacks, rebates, discounts, coupons or cheaper gift cards.

Leave a Reply